What is VRAM?
VRAM (video RAM) is the memory attached directly to a graphics card's GPU chip, where the GPU keeps the data it is working on. For AI work that means a model's weights, its KV cache and intermediate results, so a card's VRAM is the first limit on which models it can run.
VRAM is the same pool the architecture pages call GPU RAM; this page covers the practical side. Consumer cards build it from GDDR chips: the RTX 4090 has 24GB of GDDR6X and the RTX 5090 has 32GB of GDDR7. Data center cards use stacked HBM instead: the SXM H100 has 80GB of HBM3, the H200 has 141GB of HBM3e, and the AMD MI300X has 192GB of HBM3.
VRAM vs system RAM
System RAM serves the CPU and plugs into the motherboard. VRAM serves the GPU and sits on the graphics card. They hold different data and run at very different speeds.
| VRAM | System RAM | |
|---|---|---|
| Used by | The GPU | The CPU |
| Location | On the graphics card | Motherboard DIMM slots |
| Typical memory | GDDR6X, GDDR7, HBM3 | DDR5 |
| Upgrade | Fixed when you buy the card | Add or swap sticks |
The speed gap is large. A dual-channel DDR5-5600 desktop moves about 90GB/s (5,600 million transfers per second x 8 bytes x 2 channels = 89.6GB/s). The RTX 4090 moves about 1,008GB/s and the SXM H100 moves 3.35TB/s.
When a job needs more VRAM than the card has, it either fails with an out-of-memory error or spills into system RAM across the PCIe link, which is far slower than VRAM and usually makes the job crawl. Windows lists that borrowed system RAM as "Shared GPU memory", and it does not count as VRAM.
How to check how much VRAM you have
- Windows: open Task Manager (Ctrl+Shift+Esc), go to Performance, and click your GPU in the left list. "Dedicated GPU memory" is your VRAM. Ignore "Shared GPU memory" when sizing a model.
- Linux or a server: run
nvidia-smi. The Memory-Usage column shows used and total memory for each GPU, andnvidia-smi --query-gpu=name,memory.total,memory.used --format=csvprints just those fields. - Apple silicon Macs: these use unified memory, so the CPU and GPU share one pool and there is no separate VRAM figure. The memory total in About This Mac is the pool the GPU draws from.
How much VRAM AI models need
Start with the weights: parameters x bytes per parameter. 16-bit formats (FP16 and similar) use 2 bytes per parameter, 8-bit formats use 1, and 4-bit formats use about 0.5 plus a little overhead for scaling factors. Then add room for the KV cache, activations and framework overhead, which grow with context length and the number of simultaneous requests. Training and full fine-tuning need several times the weight size, since gradients and optimizer states are stored too.
The VRAM calculator on this site uses a simple rule for serving: weights x 1.2, a rule of thumb for a single request at moderate context length, not a guarantee. Using the nominal 8B and 70B parameter counts of Llama 3.1:
| Model | Precision | Weights | With 1.2x headroom | Fits on |
|---|---|---|---|---|
| Llama 3.1 8B | FP16 | 8 x 2 = 16GB | 19.2GB | One RTX 4090 (24GB) |
| Llama 3.1 70B | FP16 | 70 x 2 = 140GB | 168GB | One MI300X (192GB); two H100s (160GB) fall just short |
| Llama 3.1 70B | 4-bit | 70 x 0.5 = 35GB | 42GB | One H100 (80GB); a 32GB RTX 5090 cannot hold it |
What it means when you pick a GPU
Size the GPU from the model, not the other way round. Compute the weights, add headroom for the KV cache, and choose the fewest cards whose combined VRAM clears that total. In the table, an 8B model at FP16 fits a 24GB card with room to spare, while a 70B model at FP16 needs 168GB with headroom, more than two 80GB H100s hold.
Two traps are worth knowing. First, weights that nearly fill a card leave nothing for the KV cache. A 141GB H200 holds 140GB of 70B FP16 weights, but only about 1GB (141 - 140) is left, so long contexts or several users at once will run it out of memory. Second, once a model fits, extra VRAM does not make it faster. Speed then depends on memory bandwidth, which is what HBM provides.
The VRAM calculator does this arithmetic for a model you pick, and the GPU recommender lists which cards clear it. For a longer walkthrough, read how much VRAM you need for LLMs. Aquanode rents GPUs by the hour, with current rates on the pricing page, so you can pay for the VRAM a job needs instead of guessing high.
Building on GPUs? Aquanode runs the workload.
Deploy on H100, H200, B200, A100 and MI300X across a multi-provider marketplace, without racking your own hardware or committing to one cloud's spec sheet.
See also
GPU RAM
GPU RAM is the large off-die memory pool every Streaming Multiprocessor shares, built from slower, denser DRAM cells rather than the SRAM used in registers and cache.
HBM (High Bandwidth Memory)
HBM is stacked DRAM packaged beside a GPU die, giving data center cards several TB/s of memory bandwidth. Why LLM inference depends on it, and HBM vs GDDR.
KV Cache
A KV cache stores the key and value tensors of past tokens so an LLM never recomputes them. It grows with context length and batch size, and it eats VRAM.
nvidia-smi
nvidia-smi is the command line tool for querying and managing NVIDIA GPUs, built on the NVML management library. What it reports, what it can change, and why its text output isn't a stable interface.
Quantization
Quantization stores a model's weights, and sometimes activations, in lower-precision formats like INT8 or 4-bit, cutting VRAM use for a small accuracy cost.
Mixture of Experts (MoE)
A Mixture of Experts model routes each token to a few expert sub-networks, so compute follows active parameters but VRAM must hold every expert.