RTX 3090 for AI: 24 GB, NVLink and What Fits (2026)

Back
Team Aquanode

Team Aquanode

Sarthak Vaish

Updated OCTOBER 8, 2026Published OCTOBER 8, 2026

The RTX 3090 has 24 GB of GDDR6X on a 384-bit bus, 10,496 CUDA cores, 350 W of card power, and NVLink support, according to NVIDIA's RTX 3090 page. Its launch price was $1,499. Two years of newer cards later, it is still relevant for AI for one reason: it is the cheapest way to get 24 GB of VRAM, and the last GeForce flagship that can be bridged to a second card.

This post covers the specs, what fits in that 24 GB, the NVLink question people search for, and where the 3090 loses to the RTX 4090. It belongs to our consumer GPUs for AI guide.

TL;DR

  • Specs: 24 GB GDDR6X, 384-bit, 10,496 CUDA cores, 350 W, NVLink (two cards via a bridge).
  • VRAM is the same as the 4090. Anything that fits on a 4090 by memory fits on a 3090; the 4090 is faster, not bigger.
  • NVLink: supported on the 3090, not on the 4090. It links two cards, no more.
  • Fine-tuning: QLoRA on 7B to 14B models fits in 24 GB (computed), just as on the 4090.
  • Where it loses: a newer architecture and the 4090's higher core count. We do not publish a benchmark number for the gap.

RTX 3090 specs

Figures are from NVIDIA's RTX 3090 page and its RTX 30 Series announcement (September 1, 2020).

SpecRTX 3090RTX 4090 (for reference)
ArchitectureAmpereAda Lovelace
CUDA cores10,49616,384
Boost clock1.70 GHz2.52 GHz
Memory24 GB GDDR6X24 GB GDDR6X
Memory interface384-bit384-bit
Memory bandwidth936 GB/s at a 19.5 Gbps memory data rate on a 384-bit bus, per NVIDIA's Ampere GA102 whitepaper (not stated on the product pages)Not stated on NVIDIA's page
Power350 W card power, 750 W recommended system power450 W total graphics power, 850 W minimum PSU
NVLinkYes (SLI-Ready bridge)No
Launch priceFrom $1,499; NVIDIA's announcement gives a September 24, 2020 release date in its table and September 17 in its textFrom $1,599, October 12, 2022

Prices are 2020 and 2022 list prices. The 4090 row comes from its NVIDIA page.

Same memory size and bus width, so the bandwidth gap comes from the memory data rate, which NVIDIA's whitepapers give as 19.5 Gbps for the 3090. The compute gap is visible in the core count and clock.

Does the RTX 3090 support NVLink?

Yes. NVIDIA's spec table lists "NVIDIA NVLink (SLI-Ready)" for the 3090, and the announcement says a bridge connector links two 30 Series cards. It is a two-card link: you cannot chain three or four.

What NVLink does and does not change for LLMs:

  • Inference: you can split a model across two cards without a bridge. Frameworks such as llama.cpp and vLLM shard weights across devices over PCIe. A bridge helps when the cards must exchange data every step, as in tensor parallelism, and matters less for layer-by-layer splits. We do not state a speedup because we have not measured one.
  • Memory: two 3090s are 48 GB of total VRAM but still two separate 24 GB devices to your software. A model is split across them, not stored in one flat pool.
  • Training: gradient sync over NCCL is the traffic that benefits most from a faster link.

The 4090 dropped NVLink, so two 4090s always talk over PCIe.

What fits in 24 GB

Same arithmetic as the VRAM sizing guide: parameters times bytes per parameter, plus the KV cache, plus about 20% overhead. All numbers are computed.

ModelFP168-bit4-bitOne 3090 (24 GB)Two 3090s (48 GB)
Llama 3.1 8B16 GB8 GB4 GBAll precisionsAll precisions
Qwen3-14B (14.8B)29.6 GB14.8 GB7.4 GB8-bit, 4-bitAll precisions
Qwen3-32B (32.8B)65.6 GB32.8 GB16.4 GB4-bit8-bit, 4-bit
Llama 3.1 70B140 GB70 GB35 GB (42 with overhead)No4-bit, with little room for long context

The bottom-right cell is the reason people buy a pair. A 70B model at 4-bit is about 42 GB with overhead, so it fits in 48 GB with a few GB left for the KV cache. Push the context window and it stops fitting.

The RTX 3090 VRAM calculator lets you size a specific model.

Local LLMs on a 3090

The software story is identical to the 4090. Ollama runs a model with one command:

ollama run gemma4

llama.cpp offers llama-server with a GGUF file and layer offload via -ngl. See what is Ollama and the llama.cpp guide. Ampere does not have the FP8 tensor cores that Ada added, so FP8 checkpoints are served less natively here than on a 4090. Check the engine's docs for what it does on Ampere before choosing a format.

Image generation and fine-tuning

ComfyUI runs on the 3090, and the 24 GB is the same headroom the 4090 has for stacking models. The 3090 is simply slower per image; we do not publish a figure.

For fine-tuning, Unsloth supports Ampere cards, and its docs list the VRAM it needs per model. The computed QLoRA picture matches the 4090: an 8B base in 4-bit is about 4 GB, a 14B base about 7.4 GB, leaving the rest for activations. Full fine-tuning does not fit. See LoRA fine-tuning.

3090 vs 4090

If you already own a 3090, you do not need to upgrade to fit the same models. The 4090 buys speed and Ada features; it does not buy memory. The 3090 vs 4090 comparison and the RTX 3090 page have the side-by-side data. If you want more than 24 GB on one card, look at the RTX 5090.

Power, cooling and practical limits

NVIDIA lists 350 W card power and a 750 W recommended system power supply for the 3090, against 450 W and 850 W for the 4090. Two 3090s therefore mean 700 W of card power before the rest of the system (computed from the listed figures), which is a real constraint on a home PSU and on case airflow. The 3090 Founders Edition is a three-slot design, so check slot spacing before planning a bridged pair; NVIDIA's announcement notes the Founders Edition design shrank the NVLink and power connectors to make room for the cooler.

Memory also matters for sustained jobs. The 3090 is a consumer card with GDDR6X, and fine-tuning runs that last hours keep the memory and power budget loaded the whole time. For long runs, make sure your cooling and power headroom are real, not nominal.

Which models to start with

A sensible way to use 24 GB: run an 8B model at 8-bit for daily chat (about 8 GB of weights, computed), keep a 14B model at 4-bit (about 7.4 GB) for harder tasks, and reserve the full 24 GB for a 32B model at 4-bit (about 16.4 GB, 19.7 GB with overhead, computed). That leaves a few GB of KV cache, so keep the context window modest. Quantization is what makes all of this possible, and formats such as GGUF and AWQ are the ones you will meet most. For the full sizing rule, read how much VRAM you need for LLMs.

Run it on a cloud GPU

If you want to try 24 GB before buying, or you need a bigger card for one job, rent one. The box shows what is available now.

FAQ

Does the RTX 3090 support NVLink?

Yes, with a bridge between two cards. NVIDIA's spec table lists NVLink (SLI-Ready) for the 3090.

How much VRAM does the RTX 3090 have?

24 GB of GDDR6X on a 384-bit bus.

Is the RTX 3090 still good for AI?

For models up to roughly 32B at 4-bit on one card, yes, because its memory matches the 4090. It is slower on compute.

RTX 3090 vs 4090 for AI: which should I get?

If you need speed, the 4090. If you only need 24 GB at a lower cost, or want two NVLinked cards, the 3090. Both have 24 GB.

Can two RTX 3090s run a 70B model?

At 4-bit, yes: about 35 GB of weights, 42 GB with 20% overhead (computed), inside 48 GB. Long contexts reduce the headroom.

Sources

#consumer gpu#rtx 3090#nvlink#vram#local llm#ampere

Submit the job. Everything after that is ours.

Sign up in 60 seconds. Pay for the GPU minutes you actually use.

© 2026 Aquanode. All rights reserved.

All trademarks, logos and brand names are the property of their respective owners.