The RTX 3090 has 24 GB of GDDR6X on a 384-bit bus, 10,496 CUDA cores, 350 W of card power, and NVLink support, according to NVIDIA's RTX 3090 page. Its launch price was $1,499. Two years of newer cards later, it is still relevant for AI for one reason: it is the cheapest way to get 24 GB of VRAM, and the last GeForce flagship that can be bridged to a second card.
This post covers the specs, what fits in that 24 GB, the NVLink question people search for, and where the 3090 loses to the RTX 4090. It belongs to our consumer GPUs for AI guide.
TL;DR
- Specs: 24 GB GDDR6X, 384-bit, 10,496 CUDA cores, 350 W, NVLink (two cards via a bridge).
- VRAM is the same as the 4090. Anything that fits on a 4090 by memory fits on a 3090; the 4090 is faster, not bigger.
- NVLink: supported on the 3090, not on the 4090. It links two cards, no more.
- Fine-tuning: QLoRA on 7B to 14B models fits in 24 GB (computed), just as on the 4090.
- Where it loses: a newer architecture and the 4090's higher core count. We do not publish a benchmark number for the gap.
RTX 3090 specs
Figures are from NVIDIA's RTX 3090 page and its RTX 30 Series announcement (September 1, 2020).
| Spec | RTX 3090 | RTX 4090 (for reference) |
|---|---|---|
| Architecture | Ampere | Ada Lovelace |
| CUDA cores | 10,496 | 16,384 |
| Boost clock | 1.70 GHz | 2.52 GHz |
| Memory | 24 GB GDDR6X | 24 GB GDDR6X |
| Memory interface | 384-bit | 384-bit |
| Memory bandwidth | 936 GB/s at a 19.5 Gbps memory data rate on a 384-bit bus, per NVIDIA's Ampere GA102 whitepaper (not stated on the product pages) | Not stated on NVIDIA's page |
| Power | 350 W card power, 750 W recommended system power | 450 W total graphics power, 850 W minimum PSU |
| NVLink | Yes (SLI-Ready bridge) | No |
| Launch price | From $1,499; NVIDIA's announcement gives a September 24, 2020 release date in its table and September 17 in its text | From $1,599, October 12, 2022 |
Prices are 2020 and 2022 list prices. The 4090 row comes from its NVIDIA page.
Same memory size and bus width, so the bandwidth gap comes from the memory data rate, which NVIDIA's whitepapers give as 19.5 Gbps for the 3090. The compute gap is visible in the core count and clock.
Does the RTX 3090 support NVLink?
Yes. NVIDIA's spec table lists "NVIDIA NVLink (SLI-Ready)" for the 3090, and the announcement says a bridge connector links two 30 Series cards. It is a two-card link: you cannot chain three or four.
What NVLink does and does not change for LLMs:
- Inference: you can split a model across two cards without a bridge. Frameworks such as llama.cpp and vLLM shard weights across devices over PCIe. A bridge helps when the cards must exchange data every step, as in tensor parallelism, and matters less for layer-by-layer splits. We do not state a speedup because we have not measured one.
- Memory: two 3090s are 48 GB of total VRAM but still two separate 24 GB devices to your software. A model is split across them, not stored in one flat pool.
- Training: gradient sync over NCCL is the traffic that benefits most from a faster link.
The 4090 dropped NVLink, so two 4090s always talk over PCIe.
What fits in 24 GB
Same arithmetic as the VRAM sizing guide: parameters times bytes per parameter, plus the KV cache, plus about 20% overhead. All numbers are computed.
| Model | FP16 | 8-bit | 4-bit | One 3090 (24 GB) | Two 3090s (48 GB) |
|---|---|---|---|---|---|
| Llama 3.1 8B | 16 GB | 8 GB | 4 GB | All precisions | All precisions |
| Qwen3-14B (14.8B) | 29.6 GB | 14.8 GB | 7.4 GB | 8-bit, 4-bit | All precisions |
| Qwen3-32B (32.8B) | 65.6 GB | 32.8 GB | 16.4 GB | 4-bit | 8-bit, 4-bit |
| Llama 3.1 70B | 140 GB | 70 GB | 35 GB (42 with overhead) | No | 4-bit, with little room for long context |
The bottom-right cell is the reason people buy a pair. A 70B model at 4-bit is about 42 GB with overhead, so it fits in 48 GB with a few GB left for the KV cache. Push the context window and it stops fitting.
The RTX 3090 VRAM calculator lets you size a specific model.
Local LLMs on a 3090
The software story is identical to the 4090. Ollama runs a model with one command:
ollama run gemma4
llama.cpp offers llama-server with a GGUF file and layer offload via -ngl. See what is Ollama and the llama.cpp guide. Ampere does not have the FP8 tensor cores that Ada added, so FP8 checkpoints are served less natively here than on a 4090. Check the engine's docs for what it does on Ampere before choosing a format.
Image generation and fine-tuning
ComfyUI runs on the 3090, and the 24 GB is the same headroom the 4090 has for stacking models. The 3090 is simply slower per image; we do not publish a figure.
For fine-tuning, Unsloth supports Ampere cards, and its docs list the VRAM it needs per model. The computed QLoRA picture matches the 4090: an 8B base in 4-bit is about 4 GB, a 14B base about 7.4 GB, leaving the rest for activations. Full fine-tuning does not fit. See LoRA fine-tuning.
3090 vs 4090
If you already own a 3090, you do not need to upgrade to fit the same models. The 4090 buys speed and Ada features; it does not buy memory. The 3090 vs 4090 comparison and the RTX 3090 page have the side-by-side data. If you want more than 24 GB on one card, look at the RTX 5090.
Power, cooling and practical limits
NVIDIA lists 350 W card power and a 750 W recommended system power supply for the 3090, against 450 W and 850 W for the 4090. Two 3090s therefore mean 700 W of card power before the rest of the system (computed from the listed figures), which is a real constraint on a home PSU and on case airflow. The 3090 Founders Edition is a three-slot design, so check slot spacing before planning a bridged pair; NVIDIA's announcement notes the Founders Edition design shrank the NVLink and power connectors to make room for the cooler.
Memory also matters for sustained jobs. The 3090 is a consumer card with GDDR6X, and fine-tuning runs that last hours keep the memory and power budget loaded the whole time. For long runs, make sure your cooling and power headroom are real, not nominal.
Which models to start with
A sensible way to use 24 GB: run an 8B model at 8-bit for daily chat (about 8 GB of weights, computed), keep a 14B model at 4-bit (about 7.4 GB) for harder tasks, and reserve the full 24 GB for a 32B model at 4-bit (about 16.4 GB, 19.7 GB with overhead, computed). That leaves a few GB of KV cache, so keep the context window modest. Quantization is what makes all of this possible, and formats such as GGUF and AWQ are the ones you will meet most. For the full sizing rule, read how much VRAM you need for LLMs.
Run it on a cloud GPU
If you want to try 24 GB before buying, or you need a bigger card for one job, rent one. The box shows what is available now.
FAQ
Does the RTX 3090 support NVLink?
Yes, with a bridge between two cards. NVIDIA's spec table lists NVLink (SLI-Ready) for the 3090.
How much VRAM does the RTX 3090 have?
24 GB of GDDR6X on a 384-bit bus.
Is the RTX 3090 still good for AI?
For models up to roughly 32B at 4-bit on one card, yes, because its memory matches the 4090. It is slower on compute.
RTX 3090 vs 4090 for AI: which should I get?
If you need speed, the 4090. If you only need 24 GB at a lower cost, or want two NVLinked cards, the 3090. Both have 24 GB.
Can two RTX 3090s run a 70B model?
At 4-bit, yes: about 35 GB of weights, 42 GB with 20% overhead (computed), inside 48 GB. Long contexts reduce the headroom.