The RTX 5090 has exactly twice the memory of the RTX 5080 (32 GB against 16 GB), twice the bus width (512-bit against 256-bit) and about twice the CUDA cores (21,760 against 10,752). For AI, that memory gap is the whole story: the 5080's 16 GB caps you at 8B models at 8-bit or 14B models at 4-bit, while the 5090 reaches 32B at 4-bit.
Both are Blackwell cards with 5th-generation Tensor Cores. The 5080 is half the launch list price ($999 against $1,999), so the real question is whether your models need more than 16 GB.
TL;DR
- Buy the RTX 5080 if your models fit in 16 GB: 7B to 8B models at 8-bit, 14B at 4-bit, SDXL-class image models.
- Buy the RTX 5090 if you want 14B at 8-bit, 32B at 4-bit, or FLUX-class image models without offloading.
- Bandwidth: 960 GB/s on the 5080, 1,792 GB/s on the 5090.
- Power: 360 W against 575 W total graphics power.
- Neither runs a 70B model on one card. Rent a larger GPU for that.
Spec table
| Spec | RTX 5080 | RTX 5090 |
|---|---|---|
| Architecture | Blackwell | Blackwell |
| CUDA cores | 10,752 | 21,760 |
| Memory | 16 GB GDDR7 | 32 GB GDDR7 |
| Memory bus | 256-bit | 512-bit |
| Memory bandwidth | 960 GB/s | 1,792 GB/s |
| AI TOPS (Tensor) | 1,801 | 3,352 |
| Total graphics power | 360 W | 575 W |
| Launch list price | $999 (Jan. 30, 2025) | $1,999 (Jan. 30, 2025) |
Every row is from NVIDIA's own pages (Sources below). The two cards share a generation, so the comparison is a clean read on what the extra silicon and memory buy. The compare page shows them side by side with live rental availability.
What fits in 16 GB vs 32 GB
Computed: weights equal parameters times bytes per parameter, plus 20% overhead for KV cache and buffers (our assumption). Longer contexts or bigger batches need more.
| Model size | FP16 with overhead | INT8 with overhead | INT4 with overhead | RTX 5080 (16 GB) | RTX 5090 (32 GB) |
|---|---|---|---|---|---|
| 8B | 19.2 GB | 9.6 GB | 4.8 GB | INT8, INT4 | FP16, INT8, INT4 |
| 14B | 33.6 GB | 16.8 GB | 8.4 GB | INT4 only | INT8, INT4 |
| 32B | 76.8 GB | 38.4 GB | 19.2 GB | None | INT4 |
| 70B | 168 GB | 84 GB | 42 GB | None | None |
The 14B INT8 case (16.8 GB with overhead) misses the 5080 by under 1 GB, which is a typical example of why the cutoff matters: an 8-bit 14B model plus a modest context does not load. See how much VRAM you need for LLMs for the formula.
Speed: bandwidth and compute
Token generation is mostly bound by memory bandwidth, and the 5090's 1,792 GB/s is about 87% higher than the 5080's 960 GB/s (computed from NVIDIA's figures). The 5090 also has roughly double the AI TOPS. We do not quote tokens per second because we have not opened a benchmark from a named source for this pair. Where a model fits on both, expect the 5090 to be faster, but how much depends on the engine, the quantization and the context length.
For models that fit on both, the 5080 is not slow, it is just capacity limited. If a model fits, you give up speed. If it does not, you give up the ability to run it.
Image and video generation
FLUX.1 dev is a 12 billion parameter model per its Hugging Face card. Computed: about 24 GB at BF16 and about 12 GB at 8-bit for the main transformer alone. On the 5080 you will be running 8-bit weights with text encoders offloaded. On the 5090 you can hold BF16 weights. SDXL-class models are small enough for either card. ComfyUI and similar front ends run on both.
Fine-tuning
LoRA and QLoRA are the practical route on either card (full fine-tuning needs roughly 18 bytes per parameter with AdamW). Computed: a 4-bit 8B base is about 4 GB of frozen weights, so QLoRA on 8B fits on the 5080 with room for activations. A 4-bit 32B base is about 16 GB, which fills the 5080 on its own and leaves about 16 GB of headroom on the 5090. See LoRA fine-tuning.
How to decide in three questions
- What is the largest model you want to run, and at what precision? Size it with the table above. If the overhead-inclusive number is over 16 GB, the 5080 is out.
- How long is your context? KV cache grows with sequence length. A model that loads on the 5080 with a short context can fail with a long one, and the 5090 gives you twice the slack.
- Do you run it daily? A card you use a few hours a month is cheaper to rent than to own. A card you use all day is cheaper to own.
Power and price
The 5090's 575 W total graphics power compares with the 5080's 360 W. NVIDIA's launch list prices were $999 and $1,999 on January 30, 2025, and street prices differ. The 5080 is the easier card to power and cool. For 1,000 W power supply guidance on the 5090, see the RTX 5090 for AI guide.
Which should you pick?
- RTX 5080: you run 7B to 14B models, you care about price and power, and you accept 4-bit for 14B.
- RTX 5090: you want room for 32B at 4-bit, longer contexts, or BF16 image models.
- Something else: a 16 GB budget option is the RTX 5070 Ti. For 48 GB or more, rent a workstation or datacenter card.
Per-card pages: RTX 5080, RTX 5090, 5080 VRAM calculator, 5090 VRAM calculator. For the wider buying guide, read consumer GPUs for AI.
Rent today
Test your model on both before buying. The box shows what is available now.
FAQ
Is the RTX 5090 worth it over the 5080 for AI?
If your models need more than 16 GB, yes, because the extra memory decides what loads. If they fit in 16 GB, the 5080 does the job at half the launch list price.
How much VRAM does the RTX 5080 have?
16 GB of GDDR7 on a 256-bit bus, per NVIDIA.
Can the RTX 5080 run a 14B model?
At 4-bit, yes (about 8.4 GB with overhead, computed). At 8-bit it is just over the limit (about 16.8 GB with overhead).
How much faster is the 5090?
NVIDIA publishes about 87% more memory bandwidth and about 86% more AI TOPS for the 5090 (computed from the table). We do not quote a measured tokens-per-second gap.
Can either card run Llama 70B?
Not on one card. A 70B model at 4-bit is about 42 GB with overhead (computed), above both 16 GB and 32 GB.
Sources
- NVIDIA, GeForce RTX 5080: https://www.nvidia.com/en-us/geforce/graphics-cards/50-series/rtx-5080/
- NVIDIA, GeForce RTX 5090: https://www.nvidia.com/en-us/geforce/graphics-cards/50-series/rtx-5090/
- NVIDIA, Compare GeForce graphics cards (bandwidth, cores, TGP, AI TOPS): https://www.nvidia.com/en-us/geforce/graphics-cards/compare/
- NVIDIA, RTX 50 series launch (prices, dates): https://nvidianews.nvidia.com/news/nvidia-blackwell-geforce-rtx-50-series-opens-new-world-of-ai-computer-graphics
- FLUX.1 dev model card: https://huggingface.co/black-forest-labs/FLUX.1-dev