What is VRAM?

VRAM (video RAM) is the memory attached directly to a graphics card's GPU chip, where the GPU keeps the data it is working on. For AI work that means a model's weights, its KV cache and intermediate results, so a card's VRAM is the first limit on which models it can run.

VRAM is the same pool the architecture pages call GPU RAM; this page covers the practical side. Consumer cards build it from GDDR chips: the RTX 4090 has 24GB of GDDR6X and the RTX 5090 has 32GB of GDDR7. Data center cards use stacked HBM instead: the SXM H100 has 80GB of HBM3, the H200 has 141GB of HBM3e, and the AMD MI300X has 192GB of HBM3.

VRAM vs system RAM

System RAM serves the CPU and plugs into the motherboard. VRAM serves the GPU and sits on the graphics card. They hold different data and run at very different speeds.

VRAMSystem RAM
Used byThe GPUThe CPU
LocationOn the graphics cardMotherboard DIMM slots
Typical memoryGDDR6X, GDDR7, HBM3DDR5
UpgradeFixed when you buy the cardAdd or swap sticks

The speed gap is large. A dual-channel DDR5-5600 desktop moves about 90GB/s (5,600 million transfers per second x 8 bytes x 2 channels = 89.6GB/s). The RTX 4090 moves about 1,008GB/s and the SXM H100 moves 3.35TB/s.

When a job needs more VRAM than the card has, it either fails with an out-of-memory error or spills into system RAM across the PCIe link, which is far slower than VRAM and usually makes the job crawl. Windows lists that borrowed system RAM as "Shared GPU memory", and it does not count as VRAM.

How to check how much VRAM you have

  • Windows: open Task Manager (Ctrl+Shift+Esc), go to Performance, and click your GPU in the left list. "Dedicated GPU memory" is your VRAM. Ignore "Shared GPU memory" when sizing a model.
  • Linux or a server: run nvidia-smi. The Memory-Usage column shows used and total memory for each GPU, and nvidia-smi --query-gpu=name,memory.total,memory.used --format=csv prints just those fields.
  • Apple silicon Macs: these use unified memory, so the CPU and GPU share one pool and there is no separate VRAM figure. The memory total in About This Mac is the pool the GPU draws from.

How much VRAM AI models need

Start with the weights: parameters x bytes per parameter. 16-bit formats (FP16 and similar) use 2 bytes per parameter, 8-bit formats use 1, and 4-bit formats use about 0.5 plus a little overhead for scaling factors. Then add room for the KV cache, activations and framework overhead, which grow with context length and the number of simultaneous requests. Training and full fine-tuning need several times the weight size, since gradients and optimizer states are stored too.

The VRAM calculator on this site uses a simple rule for serving: weights x 1.2, a rule of thumb for a single request at moderate context length, not a guarantee. Using the nominal 8B and 70B parameter counts of Llama 3.1:

ModelPrecisionWeightsWith 1.2x headroomFits on
Llama 3.1 8BFP168 x 2 = 16GB19.2GBOne RTX 4090 (24GB)
Llama 3.1 70BFP1670 x 2 = 140GB168GBOne MI300X (192GB); two H100s (160GB) fall just short
Llama 3.1 70B4-bit70 x 0.5 = 35GB42GBOne H100 (80GB); a 32GB RTX 5090 cannot hold it

What it means when you pick a GPU

Size the GPU from the model, not the other way round. Compute the weights, add headroom for the KV cache, and choose the fewest cards whose combined VRAM clears that total. In the table, an 8B model at FP16 fits a 24GB card with room to spare, while a 70B model at FP16 needs 168GB with headroom, more than two 80GB H100s hold.

Two traps are worth knowing. First, weights that nearly fill a card leave nothing for the KV cache. A 141GB H200 holds 140GB of 70B FP16 weights, but only about 1GB (141 - 140) is left, so long contexts or several users at once will run it out of memory. Second, once a model fits, extra VRAM does not make it faster. Speed then depends on memory bandwidth, which is what HBM provides.

The VRAM calculator does this arithmetic for a model you pick, and the GPU recommender lists which cards clear it. For a longer walkthrough, read how much VRAM you need for LLMs. Aquanode rents GPUs by the hour, with current rates on the pricing page, so you can pay for the VRAM a job needs instead of guessing high.

Building on GPUs? Aquanode runs the workload.

Deploy on H100, H200, B200, A100 and MI300X across a multi-provider marketplace, without racking your own hardware or committing to one cloud's spec sheet.

See also

Submit the job. Everything after that is ours.

Sign up in 60 seconds. Pay for the GPU minutes you actually use.

© 2026 Aquanode. All rights reserved.

All trademarks, logos and brand names are the property of their respective owners.