llm-d is a distributed inference serving stack for production Kubernetes. It runs above model servers such as vLLM and SGLang and adds what a single engine lacks: LLM-aware routing (prefix-cache and load aware), tiered KV-cache management, prefill/decode disaggregation, wide expert parallelism and SLO-aware autoscaling. It is a CNCF sandbox project.
This guide explains the architecture, shows the official quickstart, and says when to choose it over a plain engine or over NVIDIA Dynamo. It is part of our guide to LLM inference engines.
TL;DR
- llm-d is a layer above engines, built on Kubernetes and the Gateway API Inference Extension. It does not replace vLLM or SGLang.
- Its core value is routing: send each request to the model-server pod most likely to have the prefix cached and the least load.
- It ships "well-lit paths": benchmarked recipes and Helm charts, with "Optimized Baseline" as the recommended starting point.
- You need a Kubernetes cluster with GPU nodes. For one box, a plain engine is simpler.
- The project site lists v0.7 (May 2026) as the latest release in its news section at the time of writing.
What llm-d is
The llm-d README describes it as a distributed inference serving stack for production Kubernetes deployments, running over model servers like vLLM and SGLang and adding routing, cache management and scaling. It is a CNCF sandbox project founded by Red Hat, Google Cloud, IBM Research and NVIDIA, among others.
The feature list from the README:
- Intelligent routing: prefix-cache and load-aware balancing, plus experimental predicted-latency scheduling.
- KV-cache management: tiered offloading to CPU or disk and global indexing of KV-cache state. See KV cache for the underlying idea.
- Large-model serving: prefill/decode disaggregation and wide expert parallelism (relevant to mixture-of-experts models).
- Operations: flow control for multi-tenant serving and SLO-aware autoscaling.
- Batch processing: OpenAI-compatible Batch APIs with asynchronous processing.
Architecture
The architecture docs describe three core pieces.
The Router. The entry point for requests, providing LLM-aware load balancing, queuing and policy enforcement. It has two parts: a proxy that conforms to the Gateway API Inference Extension (GAIE) and forwards requests using the ext-proc protocol, and the Endpoint Picker (EPP), the routing engine that "scores and selects model server pods based on real-time metrics, KV-cache affinity, and configured policies."
The InferencePool. A resource that groups model-server pods serving the same base model, selected by label, and acts as the discovery target for the Router. Variants are subgroups defined by pod labels, such as prefill versus decode roles.
Model servers. The engines that run the model on accelerators: vLLM and SGLang are the ones the docs cite.
Two behaviors sit on top of this:
- Prefill/decode disaggregation. A request is split into a prefill phase and a decode phase handled by specialized workers. The Router picks both endpoints and coordinates the KV-cache transfer between them. For why this helps, see TTFT and tokens per second.
- KV-cache-aware routing. Prefix-cache-aware routing uses heuristic and precise techniques to maximize cache hits, backed by a KV-cache indexer that tracks cache state across all model servers, and P2P prefix cache sharing that lets a server pull cached blocks from a peer's CPU tier instead of recomputing them. The single-engine version of prefix reuse is covered in paged attention and continuous batching.
Why routing matters for LLMs
A generic load balancer treats every request as equal, but LLM requests are not. Two requests with the same long system prompt cost far less on a replica that already holds that prefix in its KV cache than on a cold one, and a replica busy with a long prefill will make a new request wait. Round-robin ignores both facts. The Endpoint Picker reads live metrics from the pods and the cache index, so it can send the request where prefix reuse and spare capacity coincide. That is the whole pitch of the project: keep the engine you already know, and make the fleet in front of it aware of what an LLM request actually costs. The same reasoning drives the cache-aware scheduling in NVIDIA Dynamo, and the single-server version of prefix reuse is explained in paged attention and continuous batching.
Well-lit paths
Rather than a blank toolkit, llm-d publishes "well-lit paths": benchmarked recipes and Helm charts for common situations. The README names Optimized Baseline as the recommended starting point, and that is the recipe the quickstart below deploys. Other paths cover the heavier features such as prefill/decode disaggregation and wide expert parallelism. Start with the baseline, confirm it serves your model, and add features one at a time.
Hardware
The README names AMD MI300X, NVIDIA B200, and NVIDIA H100 and H200 as targets, and its release notes mention Intel XPU and Google TPU support for disaggregation. The accelerator docs in the repository hold the full list.
Quickstart
These commands are copied from the llm-d quickstart. You need kubectl, helm and a Kubernetes cluster with GPU nodes, plus a Hugging Face token.
export branch="main"
git clone https://github.com/llm-d/llm-d.git && cd llm-d && git checkout ${branch}
export REPO_ROOT=$(realpath $(git rev-parse --show-toplevel))
source ${REPO_ROOT}/guides/env.sh
export GUIDE_NAME="quickstart"
export NAMESPACE=llm-d-quickstart
Install the Gateway API Inference Extension CRDs, create the namespace and the token secret:
kubectl apply -f https://github.com/kubernetes-sigs/gateway-api-inference-extension/${GAIE_URL}/v1-manifests.yaml
kubectl create namespace ${NAMESPACE} --dry-run=client -o yaml | kubectl apply -f -
export HF_TOKEN=<your HuggingFace token>
kubectl create secret generic llm-d-hf-token \
--from-literal="HF_TOKEN=${HF_TOKEN}" \
--namespace "${NAMESPACE}" \
--dry-run=client -o yaml | kubectl apply -f -
Deploy the router in standalone mode, then the vLLM model server on NVIDIA GPUs:
helm install ${GUIDE_NAME} \
${ROUTER_STANDALONE_CHART} \
-f guides/recipes/router/base.values.yaml \
-f guides/optimized-baseline/router/optimized-baseline.values.yaml \
-n ${NAMESPACE} --version ${ROUTER_CHART_VERSION}
kubectl apply -n ${NAMESPACE} -k guides/optimized-baseline/modelserver/gpu/vllm/base/
bash ${REPO_ROOT}/guides/scripts/wait-for-pods-ready.sh -l llm-d.ai/guide=optimized-baseline -n ${NAMESPACE}
Variables such as GAIE_URL, ROUTER_STANDALONE_CHART and ROUTER_CHART_VERSION come from guides/env.sh. Then send a test request from inside the cluster:
export IP=$(kubectl get service ${GUIDE_NAME}-epp -n ${NAMESPACE} -o jsonpath='{.spec.clusterIP}')
kubectl run curl-debug --rm -it \
--image=cfmanteiga/alpine-bash-curl-jq \
--namespace="$NAMESPACE" \
--env="IP=$IP" \
-- /bin/bash
curl -X POST http://${IP}/v1/completions \
-H 'Content-Type: application/json' \
-d '{
"model": "Qwen/Qwen3-32B",
"prompt": "How are you today?"
}' | jq
Clean up:
helm uninstall ${GUIDE_NAME} -n ${NAMESPACE}
kubectl delete namespace ${NAMESPACE}
GPU and VRAM requirements
The quickstart model is Qwen3-32B. As a computed rule, weights are parameters times bytes per parameter: 32B at FP16 is about 64 GB, at FP8 about 32 GB, at INT4 about 16 GB, plus KV cache. FP16 therefore wants an 80 GB card with little room to spare, or two cards with tensor parallelism. The H100 and H200 are the natural fits; the H200's extra memory leaves more room for KV cache. Sizing details are in how much VRAM you need.
llm-d vs Dynamo vs a plain engine
| Plain engine | llm-d | Dynamo | |
|---|---|---|---|
| Scope | One server | Kubernetes serving layer | Multi-node serving framework |
| Routing | External load balancer | KV and load aware (EPP) | KV-aware Smart Router |
| Disaggregation | No | Yes | Yes |
| Engines | Itself | vLLM, SGLang | vLLM, SGLang, TensorRT-LLM |
| Best for | One model, one box | Kubernetes shops, open standards | NVIDIA-centric large clusters |
Choose llm-d when your platform is already Kubernetes and you want routing built on the Gateway API standard. Choose Dynamo when you want NVIDIA's stack including TensorRT-LLM. Choose a single vLLM server when one box is enough.
Performance
These are results reported on the llm-d site and partner blogs, each linked, and each is the reporting party's number on its own setup. We have not reproduced them.
- 3x output throughput and 2x faster TTFT with prefix-cache-aware routing versus round-robin, Llama 3.1 70B on 4 MI300X: https://llm-d.ai/blog/production-grade-llm-inference-at-scale-kserve-llm-d-vllm
- 40% lower TTFT and inter-token latency with predicted-latency scheduling versus heuristics, on NVIDIA GPUs: https://llm-d.ai/blog/predicted-latency-based-scheduling-for-llms
- About 50,000 tokens per second cluster throughput with wide expert parallelism on 16x16 B200, and 13.9x throughput with hierarchical KV offloading at 250 concurrent users versus GPU-only on 4 H100: https://llm-d.ai/blog/llm-d-v0.5-sustaining-performance-at-scale
Gains depend on how much prefix sharing and load imbalance your traffic has.
Run it on a cloud GPU
llm-d needs a Kubernetes cluster, so start by testing the model server and your model on a single GPU box, then move to the cluster recipe.
FAQ
Is llm-d an inference engine?
No. It is a serving layer above engines. vLLM or SGLang still run the model.
What is the Endpoint Picker?
The EPP is llm-d's routing engine. It scores model-server pods using real-time metrics, KV-cache affinity and configured policies, and selects where each request goes.
Does llm-d require Kubernetes?
Yes, it is built for production Kubernetes and uses the Gateway API Inference Extension.
Which GPUs does llm-d support?
The README names AMD MI300X, NVIDIA B200, H100 and H200, with Intel XPU and Google TPU mentioned for disaggregation. Check the accelerator docs for the current list.
Can I use llm-d with SGLang?
Yes, the project lists vLLM and SGLang as supported model servers.
Sources
- llm-d README (version v0.7, features, hardware, performance claims): https://github.com/llm-d/llm-d
- Quickstart: https://llm-d.ai/docs/getting-started/quickstart
- Architecture: https://llm-d.ai/docs/architecture
- Prefix-cache-aware routing result (KServe, llm-d, vLLM): https://llm-d.ai/blog/production-grade-llm-inference-at-scale-kserve-llm-d-vllm
- Predicted-latency scheduling: https://llm-d.ai/blog/predicted-latency-based-scheduling-for-llms
- llm-d v0.5 results: https://llm-d.ai/blog/llm-d-v0.5-sustaining-performance-at-scale