What is Speculative Decoding?

Speculative decoding is a way to generate text from a large language model faster without changing what it outputs. A small, fast draft model guesses the next several tokens, then the large target model checks all of those guesses in a single forward pass and keeps the ones it agrees with. When the guesses are good, one pass of the big model yields several tokens instead of one.

It works because LLM decoding leaves the GPU's arithmetic units mostly idle. Producing one token means streaming every weight of the model out of memory, so a step is limited by memory bandwidth, not by math (see arithmetic intensity and the roofline model). Checking five tokens in one pass reads the same weights as checking one, so the extra work is nearly free.

How it works

One step, in the notation of Leviathan et al. (2023):

  1. The draft model generates gamma candidate tokens one after another. Each draft step is quick because the draft model is small.
  2. The target model runs once over the prompt plus all gamma candidates and computes its own probabilities for every position at the same time.
  3. Candidates are accepted left to right. A candidate is kept if the target model's probability for it is at least the draft model's. Otherwise it is rejected with a probability set so that the final output follows exactly the target model's distribution. After the first rejection, the target model supplies a corrected token and the step ends.
  4. If every candidate is accepted, the target model's pass also yields one bonus token.

So each step produces between 1 and gamma + 1 tokens, and never fewer than plain decoding would. Because rejected guesses are corrected by sampling from an adjusted distribution, the output distribution is mathematically identical to the target model's. It is a speed-up, not an approximation, which sets it apart from quantization.

How much it saves

Let alpha be the average probability that the target accepts a draft token. The paper derives the expected tokens per step as (1 - alpha^(gamma + 1)) / (1 - alpha). If one draft step costs a fraction c of a target step, the expected wall-time speed-up is that number divided by (gamma x c + 1).

Worked example

Take gamma = 4 guesses per step.

Acceptance rate (alpha)Expected tokens per target passSpeed-up at c = 0.05Speed-up at c = 0.115
0.62.311.92x1.58x
0.83.362.80x2.30x
0.94.103.41x2.80x

The c = 0.05 column uses the ceiling the paper reports for its own draft models ("always less than 0.05"). The c = 0.115 column is our own estimate for pairing Llama 3.1 8B (8.03 billion parameters) as the draft model for Llama 3.1 70B (70.55 billion), assuming step time scales with the bytes of weights read: 8.03 / 70.55 = 0.114. Both are computed from the paper's formula, not measured, and a real serving stack adds scheduling overhead.

For reference, the paper's own measurements used a T5-XXL (11B) target at batch size 1 on a TPU-v4. Its best English-to-German result was 3.4x, with T5-small (77M) as the draft model at gamma = 7 and alpha = 0.75. The paper also showed a larger draft is not better: T5-large (800M) managed only 1.7x on the same task, because each draft step costs more and the acceptance rate gain does not pay for it. Treat these as that setup's results, not a promise for your model.

When it helps and when it does not

  • Low batch sizes help most. The method spends spare arithmetic capacity, and a lightly loaded GPU has plenty. At large batch sizes the weights are already shared across many requests and the GPU is closer to compute-bound, so the saving shrinks.
  • Predictable text helps. Code, structured output and boilerplate give high acceptance. Creative text with high sampling temperature gives low acceptance, and the extra draft work is partly wasted.
  • The draft must share the tokenizer. The draft and target need the same vocabulary, which usually means a smaller model from the same family.
  • Some variants need no separate draft model. The same verify-in-one-pass idea is used with prediction heads attached to the target model (EAGLE, multi-token prediction) or with simple n-gram matching. vLLM's documentation lists all of these and says speculative decoding targets inter-token latency under medium-to-low query rates in memory-bound workloads; see serving LLMs with vLLM and vLLM vs TensorRT-LLM vs SGLang for the engines.

What it means when you pick a GPU

Speculative decoding costs memory, not compute. You hold two sets of weights, and the draft model needs its own KV cache, so the question is whether both fit with room to spare.

Using the Llama pair above in FP8 (1 byte per parameter): the 70B target is about 70.6 GB and the 8B draft about 8.0 GB, so 78.6 GB of weights before any cache. An H100 has 80GB of HBM3, which leaves under 2 GB for caches and the CUDA context, so that pair does not work in practice. An H200 has 141GB of HBM3e, which leaves roughly 62 GB for caches, and 4.8 TB/s of memory bandwidth, so the same pair fits with room for concurrent requests. A smaller draft keeps this tighter budget: a draft in the 1B range adds only about 1 GB in FP8.

Since the technique trades spare compute for fewer weight reads, it pays off on high-bandwidth, low-concurrency interactive serving, and it does little for a batch job that already saturates the GPU. Check what fits with the VRAM calculator, look at the H200 page for its specs, and see Aquanode pricing for the cards you can rent by the hour.

Building on GPUs? Aquanode runs the workload.

Deploy on H100, H200, B200, A100 and MI300X across a multi-provider marketplace, without racking your own hardware or committing to one cloud's spec sheet.

See also

Submit the job. Everything after that is ours.

Sign up in 60 seconds. Pay for the GPU minutes you actually use.

© 2026 Aquanode. All rights reserved.

All trademarks, logos and brand names are the property of their respective owners.