RTX 5090 for AI: Specs, VRAM and What Fits (2026)

Back
Team Aquanode

Team Aquanode

Sarthak Vaish

Updated OCTOBER 8, 2026Published OCTOBER 8, 2026

The RTX 5090 is NVIDIA's flagship GeForce card for the Blackwell generation: 21,760 CUDA cores, 32 GB of GDDR7 on a 512-bit bus, 1,792 GB/s of memory bandwidth and 575 W of total graphics power. For AI the number that matters most is the 32 GB, because it decides which models load at all.

A 32B-parameter model at 4-bit fits on it with room for a long context, and a 14B model runs at 8-bit. A 70B model does not fit at any common precision.

TL;DR

  • 32 GB GDDR7 is 8 GB more than the RTX 4090's 24 GB. That is the headline upgrade for local LLMs.
  • 1,792 GB/s of bandwidth matters because LLM token generation is mostly memory-bound.
  • Fits (computed): 8B at FP16, 14B at INT8, 32B at INT4. Does not fit: 70B at INT4 (about 42 GB with overhead).
  • At 575 W it needs a serious power supply. NVIDIA lists 1000 W required system power.
  • If your model is bigger than 32 GB, rent a bigger card for the job instead of buying a second one.

RTX 5090 specs

All figures below come from NVIDIA's own pages, listed in Sources. Launch pricing is from NVIDIA's January 2025 announcement and is the launch list price, not what cards sell for today.

SpecRTX 5090RTX 4090 (for reference)
ArchitectureBlackwellAda Lovelace
CUDA cores21,76016,384
Memory32 GB GDDR724 GB GDDR6X
Memory bus512-bit384-bit
Memory bandwidth1,792 GB/snot listed on NVIDIA's 4090 page
AI TOPS (Tensor)3,352 (5th-gen Tensor Cores)1,321 (4th-gen Tensor Cores)
Total graphics power575 W450 W
Launch list price$1,999, from Jan. 30, 2025$1,599, from Oct. 12, 2022

NVIDIA's 50-series announcement also states that the Blackwell Tensor Cores support FP4, which cuts memory use for models that ship in that format. See our explainer on NVFP4 vs MXFP4 for what that means in practice.

What the specs mean for AI

VRAM is the gate. A model either fits in memory or it does not. The rest of the spec sheet only changes how fast a model that fits will run. See how much VRAM you need for LLMs for the full formula. The short version used below: weights equal parameters times bytes per parameter (FP16 is 2 bytes, INT8 is 1, INT4 is 0.5), and we add 20% for KV cache and runtime buffers. That 20% is our assumption and a long context or a large batch will need more.

Bandwidth sets generation speed. Each generated token reads roughly all the active weights from memory. A card with more bandwidth reads them faster. The 5090's 1,792 GB/s is NVIDIA's published peak, not a measured tokens-per-second figure, and we do not quote one because we have not found a benchmark from a named source that we have opened.

AI TOPS is a peak number. NVIDIA's 3,352 AI TOPS is a Tensor Core peak figure. NVIDIA does not say which precision it assumes on the product page, so treat it as a ceiling and not as a promise for FP16 workloads.

What fits in 32 GB

Computed: weights are parameters times bytes per parameter, then times 1.2 for overhead. Parameter counts are nominal sizes.

Model sizeFP16INT8INT4Fits in 32 GB?
8B16 GB (19.2 with overhead)8 GB (9.6)4 GB (4.8)All three
14B28 GB (33.6)14 GB (16.8)7 GB (8.4)INT8 and INT4. FP16 is over the limit.
32B64 GB (76.8)32 GB (38.4)16 GB (19.2)INT4 only
70B140 GB (168)70 GB (84)35 GB (42)None

Two practical notes. First, the 32B INT4 case leaves about 12 GB for KV cache, which is comfortable for single-user chat at moderate context but not for long-context batch serving. Second, quantization has a quality cost. Validate output on your own task before relying on a 4-bit model. Our glossary entries on quantization, GPTQ, AWQ and GGUF cover the formats.

You can also check any model on the RTX 5090 VRAM calculator, and the RTX 5090 page lists the card's specs and rental availability.

Running local LLMs on a 5090

The two most common local stacks are Ollama and llama.cpp. Both are open source and both use the GPU automatically on NVIDIA hardware.

Ollama's README gives the Linux install and a one-line run:

curl -fsSL https://ollama.com/install.sh | sh
ollama run gemma4

llama.cpp's README shows how to download and serve a GGUF model straight from Hugging Face:

llama serve -hf ggml-org/Qwen3.5-0.8B-GGUF

Swap in the model you want. llama.cpp's README says it supports 1.5-bit through 8-bit integer quantization, so the INT8 and INT4 rows in the table above map to real GGUF files. For a deeper walkthrough, see our Ollama guide and llama.cpp guide.

Image and video generation

Image models are far smaller than frontier LLMs, so 32 GB is generous. FLUX.1 dev is listed on its Hugging Face card as a 12 billion parameter model. Computed: 12B at BF16 is about 24 GB of weights, and at 8-bit about 12 GB. That is before the text encoders and the VAE, which ComfyUI loads alongside it, so the 24 GB BF16 figure is a floor and not a total. A 32 GB card gives you room to keep the text encoders resident instead of offloading them to system RAM. ComfyUI is the usual front end and runs on a single consumer GPU.

Fine-tuning on a 5090

Full fine-tuning needs roughly 18 bytes per parameter with mixed-precision AdamW (see the sizing guide linked above), so even an 8B model needs well past 32 GB. LoRA and QLoRA freeze the base weights, which is why they are the practical route on a consumer card. Computed: a 32B model in 4-bit is about 16 GB of frozen weights, which leaves headroom on 32 GB for adapters, optimizer state and activations at short sequence lengths. For methods, see LoRA fine-tuning and what AI model fine-tuning is.

When the 5090 is the wrong card

  • Your model is over 32 GB. A 70B model at 4-bit needs about 42 GB (computed). That needs a 48 GB class or larger card, such as the RTX PRO 6000 or a datacenter card. See the 5090 vs RTX PRO 6000 compare page.
  • You need it for a few days a month. A $1,999 launch list price is hard to justify for occasional jobs.
  • You are serving many users. Datacenter cards such as the H100 are built for that. See the H100 vs 5090 compare page.

For the full buying picture across cards, read our consumer GPUs for AI guide and the RTX 5090 vs RTX 4090 comparison.

Rent today

If you want to try a 32 GB card before buying one, or you need a bigger one for a single run, the live box below shows what is available right now.

FAQ

How much VRAM does the RTX 5090 have?

32 GB of GDDR7 on a 512-bit bus, per NVIDIA's RTX 5090 page.

Can the RTX 5090 run a 70B model?

Not on one card at common precisions. Computed: 70B at 4-bit is about 35 GB of weights, 42 GB with 20% overhead, against 32 GB of VRAM. You can offload part of the model to system RAM, which works but is much slower, or split across two cards.

Is the RTX 5090 good for AI?

For local inference of models up to about 32B at 4-bit, image generation and LoRA or QLoRA fine-tuning, yes. For larger models or multi-user serving you want more memory.

What is the RTX 5090's AI TOPS?

3,352 AI TOPS from 5th-generation Tensor Cores, per NVIDIA. It is a peak figure tied to low-precision formats.

How much power does the RTX 5090 use?

NVIDIA lists 575 W total graphics power and a 1000 W required system power on the product page.

Should I buy a 5090 or rent a GPU?

Buy if you run it most days. Rent if your use is bursty or your model needs more than 32 GB.

Sources

#consumer gpu#rtx 5090#blackwell#local llm#vram#image generation

Submit the job. Everything after that is ours.

Sign up in 60 seconds. Pay for the GPU minutes you actually use.

© 2026 Aquanode. All rights reserved.

All trademarks, logos and brand names are the property of their respective owners.