What is a KV Cache?

A KV cache is the stored copy of the key and value tensors a transformer computes for every token it has already processed, kept in GPU memory so the model does not recompute them each time it generates a new token. It exists because autoregressive decoding produces one token at a time, and every new token's attention step needs the keys and values of all the tokens before it.

That makes the cache the second big consumer of VRAM after the model weights. It is also why a model that loads fine can still run out of memory on a long prompt or a large batch: the weights are a fixed cost, but the cache keeps growing.

Why the KV cache grows

Every layer stores one key vector and one value vector per attention head for every token of every sequence in flight. The size is:

KV cache bytes = 2 x layers x kv_heads x head_dim x sequence_length x batch x bytes_per_value

The leading 2 is one tensor for keys and one for values. The cache is linear in sequence length and linear in batch size: double the context and the cache doubles, serve twice as many users at once and it doubles again. During prefill the prompt's keys and values are computed in one pass and stored. During decoding each step appends one token's worth and then reads the entire cache, which is why generation speed depends on memory bandwidth as well as capacity.

Worked example: Llama 3.1 8B

Llama 3.1 8B has 32 layers, 8 KV heads (it uses grouped-query attention, so its 32 query heads share 8 key/value heads) and a head dimension of 128. Stored in 16-bit, each value takes 2 bytes.

  1. Per token: 2 x 32 x 8 x 128 x 2 bytes = 131,072 bytes, which is 128 KiB.
  2. One 8,192-token sequence: 131,072 x 8,192 = 1,073,741,824 bytes, which is 1 GiB.
  3. A batch of 16 such sequences: 16 GiB.
  4. One sequence at the model's full 131,072-token context: 131,072 x 131,072 = 17,179,869,184 bytes, which is 16 GiB for a single user.

Larger models scale this up. Llama 3.1 70B has 80 layers, also 8 KV heads and a head dimension of 128, so the same formula gives 320 KiB per token, or 40 GiB for one 131,072-token sequence.

How to shrink it

  • Grouped-query attention (GQA). A design choice made when the model is trained: fewer key/value heads than query heads. Llama 3.1 8B keeps 8 instead of 32, so its cache is 4 times smaller than the same model with full multi-head attention (512 KiB per token).
  • KV cache quantization. Storing keys and values in 8-bit instead of 16-bit halves the cache, at some risk to output quality that you should test on your own task. See quantization.
  • Smarter memory management. vLLM uses paged attention to store the cache in fixed-size blocks, so fragmentation does not waste VRAM; our guide to serving LLMs with vLLM covers it. FlashAttention reduces the memory traffic of the attention math but does not make the cache itself smaller.

What it means when you pick a GPU

Size VRAM as weights plus KV cache plus overhead, and do the cache part from your real context length and concurrency, not from the model name alone.

Take Llama 3.1 8B at 16-bit: the weights are about 16 GB (8 billion parameters x 2 bytes). On an RTX 4090 with 24GB of GDDR6X, that leaves roughly 8 GB. At 1 GiB per 8,192-token sequence, that is room for at most 7 concurrent sequences before the CUDA context and activations take their share. An H200 has 141GB of HBM3e, so the same model leaves about 125 GB, roughly 15 times more room, which is why long-context and high-concurrency serving tends to want more VRAM before it wants more compute.

Plug your own model, context length and batch size into the VRAM calculator, or look at the H200 page for its specs. Aquanode rents GPUs by the hour, so you can size for the context length you actually serve.

Building on GPUs? Aquanode runs the workload.

Deploy on H100, H200, B200, A100 and MI300X across a multi-provider marketplace, without racking your own hardware or committing to one cloud's spec sheet.

See also

Submit the job. Everything after that is ours.

Sign up in 60 seconds. Pay for the GPU minutes you actually use.

© 2026 Aquanode. All rights reserved.

All trademarks, logos and brand names are the property of their respective owners.