The RTX 4090 has 24 GB of GDDR6X on a 384-bit bus, 16,384 CUDA cores and a 450 W total graphics power rating, according to NVIDIA's product page. For AI, the 24 GB is the number that matters: it decides which models load at all, and the 4090 is the card where 7B to 32B quantized models stop being a squeeze.
This post is the spec sheet, the VRAM arithmetic, and an honest account of what 24 GB does and does not do for local LLMs, image generation and fine-tuning. It is one card in our wider consumer GPUs for AI guide.
TL;DR
- Specs: 24 GB GDDR6X, 384-bit, 16,384 CUDA cores, 450 W, no NVLink (NVIDIA lists NVLink as "No" for this card).
- Fits (computed): 8B models at FP16, 14B at FP8 or 8-bit, and 32B-class models at 4-bit. A 70B model does not fit on one card at any common precision.
- Fine-tuning: QLoRA on 7B to 14B models is the realistic ceiling. Full fine-tuning of anything 7B or larger does not fit.
- Successor: the RTX 5090 has more memory (see our 4090 vs 5090 comparison).
- No card to buy: you can rent one by the hour when a job outgrows 24 GB or you only need it for a weekend.
RTX 4090 specs
All figures below come from NVIDIA's own RTX 4090 page and launch announcement unless marked computed.
| Spec | RTX 4090 |
|---|---|
| Architecture | Ada Lovelace |
| CUDA cores | 16,384 |
| Boost clock | 2.52 GHz (base 2.23 GHz) |
| Memory | 24 GB GDDR6X |
| Memory interface | 384-bit |
| Memory bandwidth | 1,008 GB/s at a 21 Gbps memory data rate on a 384-bit bus, per NVIDIA's Ada GPU architecture whitepaper (not stated on the product page) |
| Total graphics power | 450 W |
| Minimum PSU | 850 W |
| NVLink | No |
| Launch price | From $1,599, available October 12, 2022 (NVIDIA announcement, September 20, 2022) |
Launch price is a 2022 list price, not what the card costs today.
The missing NVLink is the one spec that shapes multi-GPU plans. Two 4090s cannot pool memory over a bridge: they are two separate 24 GB devices that frameworks split work across over PCIe. For the older card that did keep NVLink, see the RTX 3090 guide.
How much model fits in 24 GB
The rule is the one from our VRAM sizing guide: weights are parameters times bytes per parameter, then you add the KV cache and roughly 20% for runtime overhead. Everything in this table is computed with that rule, not measured.
| Model | Params | FP16 (2 B) | 8-bit (1 B) | 4-bit (0.5 B) | Fits in 24 GB? |
|---|---|---|---|---|---|
| Llama 3.1 8B | 8B | 16 GB (19.2 with overhead) | 8 GB | 4 GB | All three precisions |
| Qwen3-14B | 14.8B | 29.6 GB | 14.8 GB (17.8 with overhead) | 7.4 GB | 8-bit and 4-bit |
| Qwen3-32B | 32.8B | 65.6 GB | 32.8 GB | 16.4 GB (19.7 with overhead) | 4-bit only |
| Llama 3.1 70B | 70B | 140 GB | 70 GB | 35 GB (42 with overhead) | No |
Two caveats the table hides. First, the KV cache grows with context: the Llama 3.1 8B example in our sizing guide works out to roughly 537 MB at a 4,096-token context and about 17 GB at 128K, so a "fits" at short context can become an out-of-memory error at long context. Second, 4-bit file formats such as GGUF, AWQ and GPTQ carry some extra metadata and a few higher-precision layers, so real files run a little above the bare 0.5 bytes per parameter.
For exact numbers on this card, the RTX 4090 VRAM calculator lets you pick a model and quantization.
Running LLMs locally
Two projects cover almost every local setup on a 4090.
Ollama wraps llama.cpp in a one-command runner. From the Ollama README:
ollama run gemma4
Ollama downloads the model and starts a chat; models are tagged by size and quantization in its library. See what is Ollama for the full picture.
llama.cpp gives more control. Its server README documents llama-server with a GGUF file and -ngl for how many layers to offload to the GPU. When a model is too big for 24 GB, offloading only some layers keeps the rest in system RAM. That works, but the layers in RAM run at system-memory speed, so throughput drops sharply. We do not quote a number because it depends on your CPU and RAM. Our llama.cpp guide covers the flags.
For high-throughput serving rather than a single chat, vLLM is the usual choice, and a 24 GB card limits you to the smaller models in the table above.
Image and video generation
ComfyUI is the node-based front end most people use for Stable Diffusion and Flux-class models. Its README lists the supported models and NVIDIA GPU setup. Image models are far smaller than LLMs, so 24 GB is generous for a single image pipeline. The pressure comes from stacking: a base model, ControlNet or adapter models, a text encoder and a VAE all share the same 24 GB, and video models are heavier again. Check the model card of whatever you run for its stated VRAM need; we do not publish an images-per-second figure because we have not measured one for this post.
Fine-tuning on a 4090
Unsloth is the project most often used for single-GPU LoRA and QLoRA runs on cards like this. Its documentation lists supported models and a VRAM requirements table; read it for your exact model, since it is the project's own number and we will not restate it.
The arithmetic for why QLoRA fits (computed): a 4-bit 8B base is about 4 GB of weights, and LoRA trains only small adapter matrices, so optimizer state is tiny. The remaining 20 GB goes to activations, which scale with batch size and sequence length. A 14B base in 4-bit is about 7.4 GB, still comfortable. By contrast, full fine-tuning of an 8B model with Adam needs weights, gradients and optimizer states, which our sizing guide puts at several times the weight memory, far beyond 24 GB.
See LoRA fine-tuning and the Unsloth guide for the workflow.
When 24 GB is not enough
Rent a bigger card when you hit one of these: a 70B model at 4-bit (about 35 GB before overhead), long contexts on a 32B model, or full fine-tuning. The L40S has 48 GB and the H100 has 80 GB. For the 4090 itself, our RTX 4090 page has current availability, and the 3090 vs 4090 comparison shows the two 24 GB cards side by side.
Power and system fit
NVIDIA lists 450 W total graphics power, a three-8-pin adapter or a 450 W PCIe Gen 5 cable, and an 850 W minimum power supply (the minimum assumes a Ryzen 9 5900X build; other systems may need more). The 4090 is a large card, so check case clearance. Two 4090s need a power supply that can carry about 900 W of card power alone (computed from the listed figure) and room for two thick coolers.
Which models to start with
Practical starting points for 24 GB (all sizes computed): an 8B model at FP16 (16 GB) if you want full precision, a 14B model at 8-bit (14.8 GB) for stronger quality, and a 32B model at 4-bit (16.4 GB, 19.7 GB with overhead) when you want the largest dense model that fits. Keep the context window modest on the 32B option, since the KV cache has only a few GB left. Quantization is what makes the 32B case possible, and KV cache is what limits it.
Run it on a cloud GPU
A cloud 4090 is the same 24 GB without the hardware purchase, and the box below shows what is available right now.
FAQ
How much VRAM does the RTX 4090 have?
24 GB of GDDR6X on a 384-bit interface, per NVIDIA's product page.
Can the RTX 4090 run a 70B model?
Not on one card. A 70B model needs about 35 GB for 4-bit weights alone (computed), which is more than 24 GB before the KV cache. Two 24 GB cards, or one 48 GB card, can hold it.
Does the RTX 4090 support NVLink?
No. NVIDIA's spec table lists NVLink (SLI-Ready) as "No" for the 4090. The RTX 3090 is the last GeForce flagship that supports it.
Is the RTX 4090 good for fine-tuning?
For QLoRA and LoRA on 7B to 14B models, yes. Full fine-tuning of models in that range needs more memory than 24 GB.
What is the RTX 4090's power draw?
NVIDIA lists 450 W total graphics power and recommends an 850 W minimum power supply.