For AI work the RTX 5090 beats the RTX 4090 on every published spec: 32 GB of memory instead of 24 GB, 21,760 CUDA cores instead of 16,384, a 512-bit bus instead of 384-bit, and 3,352 AI TOPS against 1,321. The price of that is 575 W instead of 450 W and a launch list price of $1,999 instead of $1,599.
The deciding factor is almost always VRAM. If your models fit in 24 GB, the 4090 still does the job. If they sit between 24 and 32 GB, the 5090 is the card that makes them fit.
TL;DR
- Pick the RTX 5090 if you want 14B models at 8-bit with room to spare, 32B models at 4-bit with a long context, or headroom for FLUX-class image models and QLoRA on 32B.
- Pick the RTX 4090 if your models already fit in 24 GB and you do not need the extra speed.
- The memory bandwidth gap is large: 1,792 GB/s against 1,008 GB/s (about 78% more).
- The 5090 draws 125 W more at the board level (575 W against 450 W) and NVIDIA lists 1000 W of required system power for it, against 850 W for the 4090.
- Neither card has NVLink. For anything over 32 GB, rent a larger card by the hour instead of stacking two.
Spec table
| Spec | RTX 5090 | RTX 4090 |
|---|---|---|
| Architecture | Blackwell | Ada Lovelace |
| CUDA cores | 21,760 | 16,384 |
| Memory | 32 GB GDDR7 | 24 GB GDDR6X |
| Memory bus | 512-bit | 384-bit |
| Memory bandwidth | 1,792 GB/s | 1,008 GB/s |
| AI TOPS (Tensor) | 3,352 (5th-gen) | 1,321 (4th-gen) |
| Total graphics power | 575 W | 450 W |
| PCIe | Gen 5 | Gen 4 |
| NVLink | No | No |
| Launch list price | $1,999 (Jan. 30, 2025) | $1,599 (Oct. 12, 2022) |
Every row is from NVIDIA's own pages, listed at the bottom. The 4090's product page does not list memory bandwidth; its 1,008 GB/s (21 Gbps GDDR6X on a 384-bit bus) comes from NVIDIA's Ada Lovelace architecture whitepaper. Use the compare page to see both side by side with live rental availability.
What fits: both cards at FP16, FP8 and Q4
Computed: weights equal parameters times bytes per parameter (FP16 is 2, FP8 is 1, Q4 is 0.5), plus 20% for runtime buffers and a modest KV cache. The 20% is our assumption. A real 4-bit GGUF file averages a little more than 4 bits per weight, so treat the Q4 column as a lower bound. The formula is explained in how much VRAM you need for LLMs.
| Model size | FP16 (with overhead) | FP8 (with overhead) | Q4 (with overhead) | RTX 4090 (24 GB) | RTX 5090 (32 GB) |
|---|---|---|---|---|---|
| 8B | 16 GB (19.2) | 8 GB (9.6) | 4 GB (4.8) | All three | All three |
| 14B | 28 GB (33.6) | 14 GB (16.8) | 7 GB (8.4) | FP8, Q4 | FP8, Q4 |
| 32B | 64 GB (76.8) | 32 GB (38.4) | 16 GB (19.2) | Q4, tight | Q4, about 12 GB spare |
| 70B | 140 GB (168) | 70 GB (84) | 35 GB (42) | None | None |
The honest reading: the 5090 does not unlock a new model size class on its own. It gives the same class more room. A 32B model at 4-bit runs on both, but the 4090 is left with about 4.8 GB after the weights and overhead, while the 5090 keeps about 12.8 GB.
KV cache headroom at a stated context
KV cache is the part that grows with context. For a concrete case, take Llama 3.1 8B at FP16 (32 layers, 8 KV heads, head dimension 128, the figures used in our VRAM sizing guide). The per-token cost is 2 x 32 x 8 x 128 x 2 bytes = 131,072 bytes (computed).
| RTX 4090 | RTX 5090 | |
|---|---|---|
| VRAM | 24 GB | 32 GB |
| Weights at FP16 | 16 GB | 16 GB |
| Left for KV cache | 8 GB | 16 GB |
| KV cache at 32K tokens, batch 1 | 4.3 GB | 4.3 GB |
| Context that would fill the remainder | about 61K tokens | about 122K tokens |
All values are computed. They ignore the runtime buffers (CUDA kernels, framework allocations), which take part of the remainder, so real limits are lower. The pattern is what matters: the 5090 holds roughly twice the context on the same model, or twice the concurrent users at the same context.
Speed: bandwidth and Tensor Cores
Token generation reads the weights from memory for every token, so more bandwidth means faster generation for the same model. The 5090's bandwidth is about 78% above the 4090's (1,792 against 1,008 GB/s). NVIDIA's RTX 50 launch material says the 5090 outperforms the 4090 by 2X, which is NVIDIA's own claim from its own tests. We do not quote tokens per second because we have not opened a benchmark from a named source for this pair.
The Blackwell Tensor Cores add FP4 support, per NVIDIA's RTX 50 announcement. FP4 stores each weight in half a byte, so an FP4 model takes about the same memory as INT4. Whether your engine uses it depends on software support. See NVFP4 vs MXFP4 and the glossary entry on FP4.
Image and video generation
FLUX.1 dev. Its Hugging Face card lists it as a 12 billion parameter model. Computed: about 24 GB at BF16 for the main transformer alone, and about 12 GB at 8-bit, before the text encoders and the VAE. On a 4090 the BF16 transformer fills the card, so most workflows run 8-bit weights or offload parts to system RAM. The 5090's 32 GB fits the BF16 transformer with some room, though a ComfyUI workflow with several models loaded can still spill over. See the Hugging Face card, whose only VRAM note is that CPU offloading can save VRAM.
SDXL. The Stable Diffusion XL base model card lists 3 billion parameters. Computed: about 6 GB at FP16 for the base weights. Either card handles it with room for LoRAs and ControlNets, and the 4090 is not the limit. Its card also mentions CPU offloading as the VRAM-saving option.
Video: Wan 2.1. The Wan 2.1 README says "The T2V-1.3B model requires only 8.19 GB VRAM" and that it can generate a 5-second 480P video on an RTX 4090 in about 4 minutes, without quantization or other optimizations. That is the README's figure for the 4090, not a measurement of ours, and we found no equivalent published figure for the 5090. The README gives no VRAM number for its 14B models in text, only the --offload_model True option to reduce GPU memory. So the 1.3B model fits both cards easily, while the 14B models are where the extra 8 GB on the 5090 starts to matter, and we cannot say more without a published number.
Fine-tuning
Full fine-tuning is out of reach on both cards for anything past a small model, because mixed-precision AdamW needs roughly 18 bytes per parameter. LoRA and QLoRA are the route on either card. Unsloth publishes minimum VRAM figures by model size in its documentation, and it calls them the absolute minimum:
| Model | QLoRA (4-bit) | LoRA (16-bit) | 4090 (24 GB) | 5090 (32 GB) |
|---|---|---|---|---|
| 8B | 6 GB | 22 GB | QLoRA, and LoRA only barely | Both |
| 14B | 8.5 GB | 33 GB | QLoRA | QLoRA |
| 27B | 22 GB | 64 GB | QLoRA, barely | QLoRA |
| 32B | 26 GB | 76 GB | Does not fit | QLoRA |
| 70B | 41 GB | 164 GB | Does not fit | Does not fit |
Source for the Unsloth columns: its requirements page, linked in Sources. The last two columns are our reading of those numbers against each card's VRAM. Because Unsloth's numbers are minimums, a 22 GB LoRA run on a 24 GB card has almost no room for longer sequences, bigger batches or gradient checkpointing choices. So the practical gain from the 5090 is QLoRA on 32B-class models, which the 4090 cannot hold. For methods, see LoRA fine-tuning and what AI model fine-tuning is.
Power, PSU, cooling and connectors
These are NVIDIA's own figures from each product page.
| RTX 5090 | RTX 4090 | |
|---|---|---|
| Total graphics power | 575 W | 450 W |
| Required system power | 1000 W (footnote: assumes a Ryzen 9 9950X system) | 850 W |
| PSU minimum on the page | 850 W | 850 W |
| Supplementary power | 4x PCIe 8-pin cables (adapter in box) or 1x 600 W PCIe Gen 5 cable | 3x PCIe 8-pin cables (adapter in box) or a 450 W or greater PCIe Gen 5 cable |
| Size | 304 mm long, 137 mm wide (page lists 2-slot, and also asks for 3-slot clearance) | 304 mm long, 137 mm wide, 3-slot |
Three practical points. First, NVIDIA's pages call the single-cable option a "PCIe Gen 5 cable", so check that your power supply has that cable (or use the adapter) before ordering. Second, NVIDIA's 5090 page is inconsistent on slot width, so measure your case against your exact partner card. Third, the 4090 page recommends leaving clearance around the card to improve airflow, and a 575 W card will want the same or more. The 5090's page describes a dual-slot cooler with a vapor chamber, but partner cards vary.
Two cards in one machine
Neither card supports NVLink: both NVIDIA spec tables list NVLink (SLI-ready) as "No". Two cards therefore talk over PCIe, which is Gen 5 on the 5090 and Gen 4 on the 4090. In practice that means you can split a model's layers across both cards to fit a bigger model, but the cards cannot pool memory as one address space and the link between them is slower than NVLink.
Computed capacity: two 4090s give 48 GB and two 5090s give 64 GB. A 70B model at 4-bit needs about 42 GB with overhead, so it fits across either pair on paper, with less KV cache to spare on 48 GB. Two 5090s have a combined total graphics power of 1,150 W (computed), well past the single-card guidance, so the power supply, circuit and cooling become the project. We have not measured multi-card throughput, and a pair costs more than a rented single card for occasional jobs. See the glossary on tensor parallelism and NCCL for why interconnect matters when you split a model.
Who should buy which
- Buy the RTX 5090 if you regularly hit the 24 GB wall with models between 24 and 32 GB, you want QLoRA on 32B models, you need long contexts on 8B to 14B models, or you run FLUX in BF16.
- Buy the RTX 4090 (or keep it) if your workloads fit in 24 GB, you want the lower power draw, and you do not need the bandwidth. A faster card does not make a model that already fits more correct.
- Rent a datacenter GPU instead of buying either if your model needs more than 32 GB (a 70B model at 4-bit is about 42 GB with overhead, computed), if you need FP16 fine-tuning of anything over 8B, if you serve many users at once, or if your use is a few days a month. A rented H100, H200 or a workstation card like the RTX PRO 6000 gives you 80 GB or more with no power supply to buy. See the RTX PRO 6000 compare page and the H100 vs 5090 compare page.
Per-card pages: RTX 5090, RTX 4090, 5090 VRAM calculator, 4090 VRAM calculator. Single-card deep dives: RTX 5090 for AI and RTX 4090 for AI. For the wider buying guide, read consumer GPUs for AI.
Rent today
You can try both cards by the hour before deciding which one to buy. The box below shows what is available now.
FAQ
Is the RTX 5090 better than the 4090 for AI?
On every published spec, yes. The biggest practical gain is 32 GB against 24 GB of VRAM, plus about 78% more memory bandwidth (computed).
How much faster is the 5090 than the 4090?
NVIDIA's RTX 50 launch material says the 5090 outperforms the 4090 by 2X, in its own tests. We have not verified a specific tokens-per-second gap, so treat that as NVIDIA's claim.
Does the 5090 run bigger models than the 4090?
Slightly bigger or with more context. A 32B model at 4-bit fits on both, but the 5090 keeps about 12 GB free after loading it. A 14B model at FP16 (about 33.6 GB with overhead, computed) fits on neither.
What were the launch prices of the 4090 and 5090?
$1,599 on October 12, 2022 for the 4090 and $1,999 on January 30, 2025 for the 5090, both NVIDIA list prices at launch. Street prices differ.
Can I run a 70B model on either?
Not on one card. A 70B model at 4-bit is about 35 GB of weights, 42 GB with overhead (computed), which exceeds both cards. Two cards can hold it on paper, without NVLink.
Do the RTX 5090 and 4090 support NVLink?
No. NVIDIA's spec tables for both list NVLink as "No".
Sources
- NVIDIA, GeForce RTX 5090: https://www.nvidia.com/en-us/geforce/graphics-cards/50-series/rtx-5090/
- NVIDIA, GeForce RTX 4090: https://www.nvidia.com/en-us/geforce/graphics-cards/40-series/rtx-4090/
- NVIDIA, Compare GeForce graphics cards: https://www.nvidia.com/en-us/geforce/graphics-cards/compare/
- NVIDIA, RTX 50 series launch (price, date, AI TOPS): https://nvidianews.nvidia.com/news/nvidia-blackwell-geforce-rtx-50-series-opens-new-world-of-ai-computer-graphics
- NVIDIA, RTX 50 series announcements (bandwidth, FP4, 2X claim): https://www.nvidia.com/en-us/geforce/news/rtx-50-series-graphics-cards-gpu-laptop-announcements/
- NVIDIA, RTX 40 series launch (4090 price, date): https://nvidianews.nvidia.com/news/nvidia-delivers-quantum-leap-in-performance-introduces-new-era-of-neural-rendering-with-geforce-rtx-40-series
- Unsloth, requirements (VRAM by model size): https://unsloth.ai/docs/get-started/beginner-start-here/unsloth-requirements
- Wan 2.1 README: https://github.com/Wan-Video/Wan2.1
- FLUX.1 dev model card: https://huggingface.co/black-forest-labs/FLUX.1-dev
- Stable Diffusion XL base 1.0 model card: https://huggingface.co/stabilityai/stable-diffusion-xl-base-1.0
- NVIDIA Ada Lovelace GPU architecture whitepaper (RTX 4090: 21 Gbps, 1,008 GB/s): https://images.nvidia.com/aem-dam/Solutions/geforce/ada/nvidia-ada-gpu-architecture.pdf