What is a KV Cache?
A KV cache is the stored copy of the key and value tensors a transformer computes for every token it has already processed, kept in GPU memory so the model does not recompute them each time it generates a new token. It exists because autoregressive decoding produces one token at a time, and every new token's attention step needs the keys and values of all the tokens before it.
That makes the cache the second big consumer of VRAM after the model weights. It is also why a model that loads fine can still run out of memory on a long prompt or a large batch: the weights are a fixed cost, but the cache keeps growing.
Why the KV cache grows
Every layer stores one key vector and one value vector per attention head for every token of every sequence in flight. The size is:
KV cache bytes = 2 x layers x kv_heads x head_dim x sequence_length x batch x bytes_per_value
The leading 2 is one tensor for keys and one for values. The cache is linear in sequence length and linear in batch size: double the context and the cache doubles, serve twice as many users at once and it doubles again. During prefill the prompt's keys and values are computed in one pass and stored. During decoding each step appends one token's worth and then reads the entire cache, which is why generation speed depends on memory bandwidth as well as capacity.
Worked example: Llama 3.1 8B
Llama 3.1 8B has 32 layers, 8 KV heads (it uses grouped-query attention, so its 32 query heads share 8 key/value heads) and a head dimension of 128. Stored in 16-bit, each value takes 2 bytes.
- Per token: 2 x 32 x 8 x 128 x 2 bytes = 131,072 bytes, which is 128 KiB.
- One 8,192-token sequence: 131,072 x 8,192 = 1,073,741,824 bytes, which is 1 GiB.
- A batch of 16 such sequences: 16 GiB.
- One sequence at the model's full 131,072-token context: 131,072 x 131,072 = 17,179,869,184 bytes, which is 16 GiB for a single user.
Larger models scale this up. Llama 3.1 70B has 80 layers, also 8 KV heads and a head dimension of 128, so the same formula gives 320 KiB per token, or 40 GiB for one 131,072-token sequence.
How to shrink it
- Grouped-query attention (GQA). A design choice made when the model is trained: fewer key/value heads than query heads. Llama 3.1 8B keeps 8 instead of 32, so its cache is 4 times smaller than the same model with full multi-head attention (512 KiB per token).
- KV cache quantization. Storing keys and values in 8-bit instead of 16-bit halves the cache, at some risk to output quality that you should test on your own task. See quantization.
- Smarter memory management. vLLM uses paged attention to store the cache in fixed-size blocks, so fragmentation does not waste VRAM; our guide to serving LLMs with vLLM covers it. FlashAttention reduces the memory traffic of the attention math but does not make the cache itself smaller.
What it means when you pick a GPU
Size VRAM as weights plus KV cache plus overhead, and do the cache part from your real context length and concurrency, not from the model name alone.
Take Llama 3.1 8B at 16-bit: the weights are about 16 GB (8 billion parameters x 2 bytes). On an RTX 4090 with 24GB of GDDR6X, that leaves roughly 8 GB. At 1 GiB per 8,192-token sequence, that is room for at most 7 concurrent sequences before the CUDA context and activations take their share. An H200 has 141GB of HBM3e, so the same model leaves about 125 GB, roughly 15 times more room, which is why long-context and high-concurrency serving tends to want more VRAM before it wants more compute.
Plug your own model, context length and batch size into the VRAM calculator, or look at the H200 page for its specs. Aquanode rents GPUs by the hour, so you can size for the context length you actually serve.
Building on GPUs? Aquanode runs the workload.
Deploy on H100, H200, B200, A100 and MI300X across a multi-provider marketplace, without racking your own hardware or committing to one cloud's spec sheet.
See also
VRAM
VRAM is the memory attached to a GPU that holds the data it works on, and it caps which AI models fit. VRAM vs RAM, how to check yours, and how much AI needs.
FlashAttention
FlashAttention is an exact attention algorithm that tiles the computation so the full attention matrix is never written to GPU memory, cutting HBM traffic.
Quantization
Quantization stores a model's weights, and sometimes activations, in lower-precision formats like INT8 or 4-bit, cutting VRAM use for a small accuracy cost.
HBM (High Bandwidth Memory)
HBM is stacked DRAM packaged beside a GPU die, giving data center cards several TB/s of memory bandwidth. Why LLM inference depends on it, and HBM vs GDDR.
Mixture of Experts (MoE)
A Mixture of Experts model routes each token to a few expert sub-networks, so compute follows active parameters but VRAM must hold every expert.
Speculative Decoding
Speculative decoding has a small draft model guess several tokens that the large model checks in one pass, so LLM output is faster with identical results.