Inference engineering is the discipline of serving trained models in production so that they respond quickly, cost little per request and stay reliable under load. Where training produces a model once, inference engineering runs it millions of times, and the work is choosing and tuning the engine, batching, numeric precision, parallelism, hardware and scaling policy to hit a latency target at the lowest cost.
This guide defines the field, walks through the six levers with a worked example, and links to a deeper guide for each. It is the conceptual companion to our guide to LLM inference engines.
TL;DR
- Three numbers define the job: latency (time to first token and time per token), throughput (tokens per second across all users) and cost per token. Improving one usually costs another.
- The six levers are engine choice, batching and scheduling, quantization, parallelism, hardware, and autoscaling with routing.
- Order of operations: measure first, fix memory and batching, then quantize, then scale out. Disaggregation and cache-aware routing come last, when traffic justifies them.
- Every claim about speed depends on workload shape. Measure with your own prompts and concurrency.
What inference engineering is
Inference engineering is the practice of serving trained models in production while optimizing throughput, latency and cost. The work spans the engine, batching, quantization, parallelism, hardware and autoscaling, which is also how we organize this guide.
It differs from adjacent jobs:
- Machine learning research decides what the model is. Inference engineering takes the model as given.
- MLOps is the broad practice of shipping and monitoring models. Inference engineering is the performance-critical slice of it for serving.
- Fine-tuning changes the weights. See our fine-tuning overview. Inference engineering is what happens after.
If you want the basics first, start with what is AI inference and the vLLM deep dive, which walks through prefill, decode and the KV cache.
The three metrics
Everything below trades among these:
- Latency. For LLMs, time to first token (TTFT) is dominated by prefill, and the gap between later tokens is decode. Our TTFT and tokens per second post defines them precisely.
- Throughput. Total output tokens per second across all concurrent requests. It rises with concurrency until the GPU saturates, while per-request latency worsens, so always state the concurrency a number was measured at.
- Cost. Rental price divided by tokens produced. A faster configuration on the same GPU lowers cost per token even though the hourly price is unchanged. The live hourly price is in the box at the end of this post, not typed here.
Prefill is compute-bound and decode is memory-bandwidth-bound, which is why different levers help different phases. See compute-bound in the glossary.
The six levers
1. Engine choice
The inference engine schedules requests, manages the KV cache and runs the kernels. The main open options are vLLM, SGLang and TensorRT-LLM, and each has packaging and orchestration layers around it. Start with our comparison vLLM vs TensorRT-LLM vs SGLang, then go deep on one:
- vLLM install and Docker
- SGLang
- TensorRT-LLM
- NVIDIA NIM for packaged NVIDIA containers
- Triton Inference Server for mixed-model serving
- FlashInfer, the attention kernel library several engines use
Smaller deployments on a workstation GPU may prefer Ollama or llama.cpp.
2. Batching and scheduling
A GPU is wasteful on one request at a time. Static batching waits for the slowest request in a batch; continuous batching lets finished requests leave and new ones join at every step. The vLLM paper (PagedAttention) reports that managing the KV cache in small blocks and batching continuously improved throughput by 2 to 4 times versus FasterTransformer and Orca at the same latency (the authors' figure on their benchmarks). Chunked prefill and prefix caching extend the idea. All of it is covered in paged attention and continuous batching.
3. Quantization
Storing weights (and sometimes the KV cache) in fewer bits shrinks memory and the bytes moved per token, which usually speeds decode. Computed example: a 70B-parameter model needs about 140 GB for weights at 16-bit, about 70 GB at 8-bit and about 35 GB at 4-bit (parameters times bytes per parameter, before KV cache). That is the difference between several GPUs and one. The trade is a possible quality drop, so evaluate on your own task. Read the glossary entries for quantization, FP8, FP4, AWQ and GPTQ, then NVFP4 vs MXFP4 and Transformer Engine FP8.
4. Parallelism
When a model or its traffic outgrows one GPU you split the work: tensor parallelism slices each layer across GPUs, pipeline parallelism places different layers on different devices, data parallelism replicates the model, and expert parallelism spreads mixture-of-experts experts. Tensor parallelism communicates constantly, so it wants fast links inside one node. Another lever in this family is speculative decoding, where a draft model proposes tokens and the main model verifies them in one pass.
Disaggregation is parallelism by phase: prefill and decode run on separate pools. The DistServe paper reports that doing so allowed serving up to 7.4 times more requests, or meeting up to 12.6 times tighter latency targets, than the systems it compared against (the authors' figures on their workloads). Production stacks that implement it are covered in NVIDIA Dynamo and llm-d.
5. Hardware
Hardware sets the ceiling: memory capacity decides what fits, memory bandwidth bounds decode speed, and compute bounds prefill. The H200 has more memory than the H100, which buys KV cache room and fewer GPUs per model, while the B200 brings newer low-precision formats. Compare them in H100 vs H200, pick cards in best GPU for LLM inference, and see the datacenter GPU overview. For sizing, use the VRAM calculator and how much VRAM you need.
6. Autoscaling and routing
A single replica is not a production system. Traffic changes, and requests that share a prefix are cheaper when they land on a replica that already holds the cache. Cache-aware routers and latency-driven autoscalers handle this: Dynamo's Planner scales prefill and decode pools to a latency target, and llm-d's Endpoint Picker scores pods on load and KV-cache affinity. Both projects publish gains from these techniques, with the caveat that they depend on how much prefix sharing and burstiness your traffic has.
A worked example: where the memory goes
This is a computed illustration with assumed numbers, not a benchmark. Take a dense model with 70 billion parameters, 80 layers, 8 key-value heads and a head dimension of 128, served in 16-bit.
- Weights: 70 billion times 2 bytes is about 140 GB.
- KV cache per token: 2 (key and value) times 80 layers times 8 heads times 128 dimensions times 2 bytes, which is 327,680 bytes, about 320 KiB.
- One 32,768-token sequence: 327,680 times 32,768 is about 10.7 GB (10 GiB).
So a handful of long conversations already adds tens of gigabytes on top of the weights. Three consequences follow, and each maps to a lever:
- Weights alone exceed one 80 GB GPU, so you need parallelism, quantization or a bigger-memory card.
- FP8 weights (about 70 GB) plus a quantized KV cache change how many sequences fit, which is a throughput decision.
- If many users share a long system prompt, prefix caching and cache-aware routing avoid recomputing and re-storing it.
A practical order of operations
- Define the target. Pick a TTFT and per-token latency budget and a realistic concurrency.
- Measure a baseline with an off-the-shelf engine on one GPU class, after a few warmup requests.
- Fix memory first. Cap context length to what you need, size the KV cache, enable prefix caching if prompts share a prefix.
- Quantize and re-check quality on your own evaluation set.
- Add parallelism only when the model does not fit or one replica cannot meet the latency target.
- Scale out with replicas, then routing and autoscaling.
- Adopt disaggregation when long prefills visibly stall decode for other users.
Aquanode manages and optimizes GPUs for training and inference workloads, and the box below shows what is available to rent today for these experiments.
Run it on a cloud GPU
Test the engine and quantization levers on a single card first, then step up to a multi-GPU node for parallelism experiments.
FAQ
What does an inference engineer do?
They make model serving faster, cheaper and more reliable: choosing the engine, tuning batching and memory, quantizing, picking parallelism and hardware, and building routing and autoscaling.
What is the difference between inference engineering and MLOps?
MLOps covers the whole lifecycle of shipping and monitoring models. Inference engineering is the performance and cost work specific to serving.
What is the biggest lever for LLM serving cost?
It depends on the workload, so measure. Batching and memory management usually come first, then quantization, because both raise how many tokens each GPU produces. Neither is a substitute for measuring your own traffic.
How does inference engineering differ for training?
Training moves gradients and optimizer state and is throughput-oriented. Inference needs only the forward pass, so memory is dominated by weights and KV cache, and latency matters.
Which inference engine should I start with?
For most teams, vLLM or SGLang on a single node, then compare against TensorRT-LLM if you are NVIDIA-only and need more performance. See the engine comparison.
Sources
- Efficient Memory Management for LLM Serving with PagedAttention (2-4x throughput claim): https://arxiv.org/abs/2309.06180
- DistServe: Disaggregating Prefill and Decoding (7.4x requests, 12.6x tighter SLO): https://arxiv.org/abs/2401.09670
- NVIDIA Dynamo README: https://github.com/ai-dynamo/dynamo
- NVIDIA Dynamo introduction (Planner, Smart Router, KV cache manager, NIXL): https://developer.nvidia.com/blog/introducing-nvidia-dynamo-a-low-latency-distributed-inference-framework-for-scaling-reasoning-ai-models/
- llm-d architecture (Endpoint Picker, KV-cache-aware routing): https://llm-d.ai/docs/architecture