TTFT and Tokens Per Second: LLM Latency Metrics (2026)

Back
Team Aquanode

Team Aquanode

Sarthak Vaish

Updated OCTOBER 8, 2026Published OCTOBER 8, 2026

TTFT (time to first token) is the delay between sending a request and receiving the first token of the answer; it is dominated by queueing and prompt processing (prefill). Tokens per second measures generation speed after that point, either for one user (the reciprocal of inter-token latency) or summed across every user the server is handling (system throughput). They pull in opposite directions, so serving is the job of choosing which one to protect.

This post defines each metric the way the engines do, shows the arithmetic that sets a floor for each on an H100, and walks through measuring them. It is part of our guide to LLM inference engines.

TL;DR

  • TTFT is a prefill and queueing number. It rises with prompt length and with how busy the server is. Prefix caching and chunked prefill are the main levers.
  • Inter-token latency (ITL) and TPOT are decode numbers. Decode is limited by memory bandwidth at small batch sizes, so quantization and fewer bytes per token help most.
  • Tokens per second means two different things. Per-user speed is about 1 divided by ITL. System throughput is total output tokens over wall time. Always say which one.
  • Never quote a number without the concurrency and prompt length it was measured at. A single-request figure and a saturated-server figure can differ by an order of magnitude.
  • Verdict: report TTFT and ITL at p50 and p99 for your real traffic shape, plus system throughput at the concurrency you plan to run. Use vllm bench serve or your engine's equivalent.

The definitions

Different tools use slightly different words, so pin them down first. NVIDIA's LLM benchmarking documentation and the vLLM benchmark docs agree on the core ones:

MetricDefinitionMostly governed by
TTFTTime from submitting a request to receiving the first tokenQueueing, prefill (prompt length)
E2E latencyTime from submitting a request to receiving the full response; TTFT plus the generation timeTTFT, ITL, output length
ITLAverage time between consecutive output tokens, excluding TTFTDecode speed
TPOTTime per output token; NVIDIA treats it as the same quantity as ITLDecode speed
TPS (system)Total output tokens divided by the total time windowBatching, hardware
TPS per userOutput tokens of one request divided by its E2E latencyITL, TTFT
RPSCompleted requests per secondEverything above

The formulas NVIDIA gives (in code form because of the symbols):

e2e_latency = TTFT + generation_time
ITL         = (e2e_latency - TTFT) / (total_output_tokens - 1)
TPS         = total_output_tokens / (end_time - start_time)

A naming trap. vLLM reports TPOT per request as (end-to-end latency minus TTFT) divided by (output tokens minus 1), then aggregates it across requests. It reports ITL separately, as the gaps between consecutive streamed outputs pooled across all requests. The two usually sit close together, but the mean ITL is not weighted by request length the way a per-request TPOT average is. When comparing numbers from two tools, check which one they print.

Worked example: what the numbers imply (computed)

Say a chat request streams 300 output tokens with a TTFT of 0.4 seconds and an ITL of 25 ms.

  • Generation time: 299 gaps x 0.025 s = 7.475 s.
  • E2E latency: 0.4 + 7.475 = 7.875 s.
  • Per-user TPS by NVIDIA's definition: 300 / 7.875 = 38.1 tokens per second, close to 1 / 0.025 = 40.

Notice how per-user TPS sits below 1 divided by ITL because TTFT is included. For short answers the gap is large, for long answers it vanishes. NVIDIA's docs make the same point: per-user TPS approaches 1 over ITL as the output grows.

What sets TTFT

TTFT has three parts: time waiting in the queue, time for prefill, and the first decode step.

Prefill is compute-bound. All prompt tokens go through the model in one parallel pass, so it is limited by tensor core throughput. A common rule of thumb is about 2 FLOPs per parameter per token (it ignores the attention term, which grows with context).

Computed floor for an 8B model with a 2,000-token prompt on an H100:

  • Work: 2 x 8 x 10^9 x 2,000 = 3.2 x 10^13 FLOPs, or 32 TFLOP.
  • At the 989 TFLOPS dense FP16 matmul peak quoted in the FlashAttention-3 write-up, that is about 32 ms at 100% utilization.

Real engines will not hit peak, so treat 32 ms as a lower bound, not a prediction. The useful lesson is the scaling: TTFT grows roughly linearly with prompt length, and the queue term often matters more than the compute term once the server is busy.

What moves TTFT in practice:

  • Prefix caching. If the prompt shares a prefix with an earlier request, the engine reuses its KV blocks and only prefills the rest. See PagedAttention and continuous batching for how vLLM and SGLang do it.
  • Chunked prefill and batch token budget. In vLLM, a higher max_num_batched_tokens gives better TTFT and a lower one gives better ITL, per the vLLM tuning docs. That is the TTFT versus ITL trade in one flag.
  • Queueing. When arrivals exceed what the server can run, requests wait. TTFT then climbs even though prefill itself is unchanged. This is why TTFT at the 99th percentile is the number to watch.
  • Attention kernel. Faster prefill kernels, such as FlashAttention and FlashInfer, cut the compute term.
  • Model and precision. Fewer FLOPs per token (a smaller model, or FP8 on hardware that runs it faster) shortens prefill.

What sets tokens per second

Decode is memory-bound at small batch sizes. Each decode step must read the whole set of weights from GPU memory to produce one token per request, and reads the KV cache for each request's context as well.

Computed ceiling for the same 8B model at FP16 (about 16 GB of weights) on an H100 with 3.35 TB/s of memory bandwidth (NVIDIA's H100 page):

  • Time to stream the weights once: 16 / 3,350 = 4.8 ms.
  • Batch of one: at most about 1 / 0.0048 = 209 tokens per second (ignoring KV cache reads, which only lower it).

Now batch more requests. A decode step reads the weights once regardless of how many sequences ride along, so throughput scales with batch size until compute becomes the limit. Computed: dividing 989 TFLOPS by 3.35 TB/s gives about 295 FLOPs per byte. At FP16, a batch of B sequences does about B FLOPs per weight byte (2 FLOPs per 2-byte parameter, per sequence), so the crossover sits near a batch of roughly 300. Below it you are bandwidth-bound and each extra request is nearly free in step time. Above it you are compute-bound. Real systems cross over earlier because KV cache reads add traffic, but the shape holds.

This is why engines batch aggressively, and why the same server shows high system throughput but modest per-user speed under load: you gave up ITL to buy aggregate tokens per second.

What moves tokens per second:

  • Fewer bytes per token. Quantization (INT8, INT4, FP8) shrinks the weights streamed per step. Quality is the cost.
  • Higher memory bandwidth. This is the main reason datacenter parts such as the H200 matter for decode-heavy work.
  • Speculative decoding. A draft model proposes tokens that the main model verifies in one pass, so you can emit several tokens per step. See speculative decoding. The gain depends heavily on the workload.
  • Tensor parallelism. Spreading the weights over more GPUs adds aggregate bandwidth, at the cost of communication.
  • A bigger KV cache budget. More resident sequences means a bigger batch. See how much VRAM you need for LLMs.

The trade-off in one picture

Plot per-user speed against system throughput as concurrency rises: system throughput climbs and flattens, per-user speed falls. TTFT tends to rise sharply near saturation as the queue builds. A good operating point is the highest concurrency that still meets your latency target for TTFT and ITL, not the peak of the throughput curve.

How to measure it with vLLM

vllm bench serve sends traffic to a running server and prints TTFT, TPOT and ITL as mean, median and p99. This is the example from the vLLM benchmarking docs:

vllm bench serve \
  --backend vllm \
  --model NousResearch/Hermes-3-Llama-3.1-8B \
  --endpoint /v1/completions \
  --dataset-name sharegpt \
  --dataset-path <your data path>/ShareGPT_V3_unfiltered_cleaned_split.json \
  --num-prompts 10

The flags that matter for a useful test:

  • --request-rate sets target requests per second. The default is inf, which sends everything at once; a finite value uses a Poisson arrival process by default.
  • --max-concurrency caps requests in flight. The docs give --request-rate=inf --max-concurrency=<limit> as the pattern for measuring maximum throughput under backpressure.
  • --num-prompts sets the sample size. Ten prompts is a smoke test, not a benchmark.
  • --dataset-name picks the workload; the docs list sharegpt, random-mm, prefix_repetition and others.

Measurement rules that save you from fooling yourself:

  1. Warm up first. The first requests pay for cache allocation and kernel compilation.
  2. Use a prompt and output length distribution close to production. A benchmark with 128-token prompts says nothing about a RAG service with 6,000-token prompts.
  3. Sweep concurrency and plot the curve instead of reporting one point.
  4. Report p99 beside the median. Tail latency is what users complain about.
  5. Count tokens with the model's own tokenizer, because token counts differ across model families.
  6. Never compare a number from another engine's blog post to yours unless the model, precision, hardware, prompt length and concurrency match.

For where these numbers come from at each pipeline stage, serving LLMs with vLLM walks through prefill and decode, and vLLM vs TensorRT-LLM vs SGLang compares engines on them.

Run it on a cloud GPU

To measure your own TTFT and ITL, run the same benchmark on two different cards and compare the curves. The L40S and the H100 differ mainly in memory bandwidth and tensor throughput, which map directly onto the decode and prefill terms above.

FAQ

What is a good TTFT?

There is no universal number. It depends on prompt length and the product. Interactive chat usually wants a first token fast enough to feel instant, while a batch summarization job does not care. Set a target from your use case and measure p99, not the average.

Are ITL and TPOT the same thing?

NVIDIA's documentation treats them as the same quantity. vLLM reports both: TPOT per request, ITL pooled across token gaps. They are close but not identical, so check what your tool prints.

Why is my tokens per second lower under load?

Because per-user speed falls as the batch grows. Each decode step takes longer when more sequences ride in it, so every user waits longer between tokens, while the sum across users goes up.

Does a longer prompt change tokens per second?

Prompt length mainly changes TTFT. It also grows the KV cache each decode step must read, which slows ITL a little and reduces how many requests fit at once.

Does quantization improve TTFT?

It can, because it reduces prefill work on hardware that runs lower precision faster, but the larger effect is on decode, where fewer bytes per weight mean a faster step. Measure both on your model.

How many tokens per second does an H100 do?

It depends on the model, precision, batch and prompt length, so no single number is honest. We do not quote one. Run the benchmark above on the model you plan to serve.

Sources

#llm inference#inference engines#ttft#tokens per second#latency#benchmarking

Submit the job. Everything after that is ours.

Sign up in 60 seconds. Pay for the GPU minutes you actually use.

© 2026 Aquanode. All rights reserved.

All trademarks, logos and brand names are the property of their respective owners.