What is Speculative Decoding?
Speculative decoding is a way to generate text from a large language model faster without changing what it outputs. A small, fast draft model guesses the next several tokens, then the large target model checks all of those guesses in a single forward pass and keeps the ones it agrees with. When the guesses are good, one pass of the big model yields several tokens instead of one.
It works because LLM decoding leaves the GPU's arithmetic units mostly idle. Producing one token means streaming every weight of the model out of memory, so a step is limited by memory bandwidth, not by math (see arithmetic intensity and the roofline model). Checking five tokens in one pass reads the same weights as checking one, so the extra work is nearly free.
How it works
One step, in the notation of Leviathan et al. (2023):
- The draft model generates gamma candidate tokens one after another. Each draft step is quick because the draft model is small.
- The target model runs once over the prompt plus all gamma candidates and computes its own probabilities for every position at the same time.
- Candidates are accepted left to right. A candidate is kept if the target model's probability for it is at least the draft model's. Otherwise it is rejected with a probability set so that the final output follows exactly the target model's distribution. After the first rejection, the target model supplies a corrected token and the step ends.
- If every candidate is accepted, the target model's pass also yields one bonus token.
So each step produces between 1 and gamma + 1 tokens, and never fewer than plain decoding would. Because rejected guesses are corrected by sampling from an adjusted distribution, the output distribution is mathematically identical to the target model's. It is a speed-up, not an approximation, which sets it apart from quantization.
How much it saves
Let alpha be the average probability that the target accepts a draft token. The paper derives the expected tokens per step as (1 - alpha^(gamma + 1)) / (1 - alpha). If one draft step costs a fraction c of a target step, the expected wall-time speed-up is that number divided by (gamma x c + 1).
Worked example
Take gamma = 4 guesses per step.
| Acceptance rate (alpha) | Expected tokens per target pass | Speed-up at c = 0.05 | Speed-up at c = 0.115 |
|---|---|---|---|
| 0.6 | 2.31 | 1.92x | 1.58x |
| 0.8 | 3.36 | 2.80x | 2.30x |
| 0.9 | 4.10 | 3.41x | 2.80x |
The c = 0.05 column uses the ceiling the paper reports for its own draft models ("always less than 0.05"). The c = 0.115 column is our own estimate for pairing Llama 3.1 8B (8.03 billion parameters) as the draft model for Llama 3.1 70B (70.55 billion), assuming step time scales with the bytes of weights read: 8.03 / 70.55 = 0.114. Both are computed from the paper's formula, not measured, and a real serving stack adds scheduling overhead.
For reference, the paper's own measurements used a T5-XXL (11B) target at batch size 1 on a TPU-v4. Its best English-to-German result was 3.4x, with T5-small (77M) as the draft model at gamma = 7 and alpha = 0.75. The paper also showed a larger draft is not better: T5-large (800M) managed only 1.7x on the same task, because each draft step costs more and the acceptance rate gain does not pay for it. Treat these as that setup's results, not a promise for your model.
When it helps and when it does not
- Low batch sizes help most. The method spends spare arithmetic capacity, and a lightly loaded GPU has plenty. At large batch sizes the weights are already shared across many requests and the GPU is closer to compute-bound, so the saving shrinks.
- Predictable text helps. Code, structured output and boilerplate give high acceptance. Creative text with high sampling temperature gives low acceptance, and the extra draft work is partly wasted.
- The draft must share the tokenizer. The draft and target need the same vocabulary, which usually means a smaller model from the same family.
- Some variants need no separate draft model. The same verify-in-one-pass idea is used with prediction heads attached to the target model (EAGLE, multi-token prediction) or with simple n-gram matching. vLLM's documentation lists all of these and says speculative decoding targets inter-token latency under medium-to-low query rates in memory-bound workloads; see serving LLMs with vLLM and vLLM vs TensorRT-LLM vs SGLang for the engines.
What it means when you pick a GPU
Speculative decoding costs memory, not compute. You hold two sets of weights, and the draft model needs its own KV cache, so the question is whether both fit with room to spare.
Using the Llama pair above in FP8 (1 byte per parameter): the 70B target is about 70.6 GB and the 8B draft about 8.0 GB, so 78.6 GB of weights before any cache. An H100 has 80GB of HBM3, which leaves under 2 GB for caches and the CUDA context, so that pair does not work in practice. An H200 has 141GB of HBM3e, which leaves roughly 62 GB for caches, and 4.8 TB/s of memory bandwidth, so the same pair fits with room for concurrent requests. A smaller draft keeps this tighter budget: a draft in the 1B range adds only about 1 GB in FP8.
Since the technique trades spare compute for fewer weight reads, it pays off on high-bandwidth, low-concurrency interactive serving, and it does little for a batch job that already saturates the GPU. Check what fits with the VRAM calculator, look at the H200 page for its specs, and see Aquanode pricing for the cards you can rent by the hour.
Building on GPUs? Aquanode runs the workload.
Deploy on H100, H200, B200, A100 and MI300X across a multi-provider marketplace, without racking your own hardware or committing to one cloud's spec sheet.
See also
KV Cache
A KV cache stores the key and value tensors of past tokens so an LLM never recomputes them. It grows with context length and batch size, and it eats VRAM.
Quantization
Quantization stores a model's weights, and sometimes activations, in lower-precision formats like INT8 or 4-bit, cutting VRAM use for a small accuracy cost.
Roofline Model
The roofline model plots a kernel's arithmetic intensity against two hardware ceilings, memory bandwidth and arithmetic bandwidth, to show at a glance whether it's compute-bound or memory-bound. Where it came from and why GPUs need it.
Arithmetic Intensity
Arithmetic intensity is the ratio of compute operations to bytes moved in a kernel. Why it decides whether a workload is compute-bound or memory-bound, and how tricks like recomputation trade memory traffic for extra FLOPs.
Compute-bound
A compute-bound kernel is limited by arithmetic throughput rather than memory bandwidth. When LLM inference hits this regime, and a back-of-the-envelope estimate for the batch size it takes to get there.
HBM (High Bandwidth Memory)
HBM is stacked DRAM packaged beside a GPU die, giving data center cards several TB/s of memory bandwidth. Why LLM inference depends on it, and HBM vs GDDR.