NVIDIA Dynamo is an open-source, datacenter-scale inference framework that sits above engines such as vLLM, SGLang and TensorRT-LLM and turns them into a coordinated multi-node serving system. Its main ideas are splitting prefill and decode onto separate GPU pools, routing requests to the worker that already holds their KV cache, tiering that cache across memory and storage, and autoscaling to a latency target.
This guide explains each piece, shows the project's own quickstart commands, and says when Dynamo is worth the complexity. It is part of our guide to LLM inference engines.
TL;DR
- Dynamo is not another engine. It orchestrates engines: vLLM, SGLang and TensorRT-LLM are the supported backends.
- Its gains come from system-level moves: disaggregated prefill and decode, KV-aware routing, KV offload (KVBM) and SLA-driven autoscaling (Planner).
- It pays off for multi-GPU, multi-node, high-traffic serving, especially long prompts and reasoning models. For one model on one box, plain vLLM is simpler.
- The docs describe version 1.5.1 as current at the time of writing; the README's container tags reference 1.5.0.
What Dynamo is
The Dynamo README calls it an open-source, datacenter-scale inference stack built in Rust for performance and Python for extensibility. NVIDIA introduced it as a "high-throughput, low-latency open-source inference serving framework" for generative AI and reasoning models. The key design choice is that it does not replace the engine. Your vLLM, SGLang or TensorRT-LLM workers still run the model. Dynamo handles the frontend, routing, scheduling between workers, cache movement and scaling.
That puts it in the same family as llm-d (a Kubernetes-native project with a similar goal) and complements packaged options like NVIDIA NIM.
How it works
Disaggregated prefill and decode
An LLM request has two phases with different bottlenecks. Prefill processes the whole prompt in one parallel pass and is compute-bound. Decode generates one token at a time and is limited by memory bandwidth. When both run on the same GPUs, long prefills stall decoding for other users. Dynamo's README describes the fix: it "separates prefill and decode into independently scaled GPU pools, so each phase runs on hardware suited to its workload." The research behind the idea is the DistServe paper, which reports that disaggregating the two phases lets a system serve up to 7.4x more requests, or meet up to 12.6x tighter latency targets, than the systems it compared against (the authors' figures, on their workloads and baselines). See TTFT and tokens per second for the metrics involved and paged attention and continuous batching for the single-engine techniques Dynamo builds on.
The four core components
From NVIDIA's introduction post:
- Planner. Monitors GPU capacity metrics in the cluster and decides whether to use disaggregated serving for a request, or whether to add GPUs to the prefill or decode phase. The README describes it as an SLA-driven autoscaler that profiles workloads and right-sizes pools.
- Smart Router. Hashes incoming requests into a radix tree to track where cached KV data lives, then routes each request to the worker with the best cache overlap while keeping load balanced. This avoids recomputing prefixes that some worker already holds.
- Distributed KV Cache Manager (KVBM). Keeps hot KV cache in GPU memory and moves colder blocks to host memory, local SSDs or networked storage, extending effective context beyond GPU memory. The README's support table marks KVBM as supported for vLLM and SGLang and in progress for TensorRT-LLM.
- NIXL. The NVIDIA Inference Transfer Library, a point-to-point communication library with one API for moving inference data between memory and storage types, selecting the best connection backend automatically. It is what lets a decode worker receive the KV cache a prefill worker built. For GPU-to-GPU transfer at this scale, interconnect matters: see our notes on NCCL and KV cache.
The README lists further components: ModelExpress (weight streaming between GPUs), Grove (Kubernetes topology-aware scheduling), AISimulate (offline configuration search) and fault tolerance with request migration.
Supported backends
SGLang, TensorRT-LLM and vLLM. Per the README table, disaggregated serving, KV-aware routing, the Planner, multimodal support and tool calling are supported on all three.
Quickstart
These commands are copied from the Dynamo README (container tags reference release 1.5.0). You need a GPU host with Docker and the NVIDIA Container Toolkit.
Option A: container, vLLM backend
docker run --gpus all --network host --rm -it nvcr.io/nvidia/ai-dynamo/vllm-runtime:1.5.0
python3 -m dynamo.frontend --http-port 8000 --discovery-backend file > /dev/null 2>&1 &
python3 -m dynamo.vllm --model Qwen/Qwen3-0.6B --discovery-backend file &
The first command starts the container; run the next two inside it. The frontend serves an OpenAI-compatible API on port 8000.
Option A: container, SGLang backend
docker run --gpus all --network host --rm -it nvcr.io/nvidia/ai-dynamo/sglang-runtime:1.5.0
python3 -m dynamo.frontend --http-port 8000 --discovery-backend file > /dev/null 2>&1 &
python3 -m dynamo.sglang --model-path Qwen/Qwen3-0.6B --discovery-backend file &
Option B: PyPI
Install uv first, then the package for your backend:
curl -LsSf https://astral.sh/uv/install.sh | sh
uv pip install --prerelease=allow "ai-dynamo[vllm]"
uv pip install --prerelease=allow "ai-dynamo[sglang]"
TensorRT-LLM via PyPI needs pip with --extra-index-url https://pypi.nvidia.com, and the README also lists a tensorrtllm-runtime:1.5.0 container. Once a worker is up, send a normal OpenAI-style request to http://localhost:8000/v1/chat/completions (see serving LLMs with vLLM for the request shape).
For disaggregated serving across workers and for Kubernetes deployment, follow the version-specific pages in the Dynamo documentation at docs.nvidia.com/dynamo; the exact launch recipes change between releases, so copy them from the docs for the version you run rather than from a blog post.
GPU and VRAM requirements
Dynamo adds no model memory of its own; each worker needs what the engine needs. As a computed rule, weights are parameters times bytes per parameter: a 70B model is about 140 GB at FP16 or about 70 GB at FP8, so it needs two or more 80 GB GPUs per worker before KV cache. Disaggregation multiplies this, because prefill and decode each hold a full copy of the weights (or a tensor-parallel slice). Plan for at least two workers per model. The H100 and H200 are the common choices, and the B200 is the NVIDIA part Dynamo's headline numbers use. See how much VRAM you need for the sizing math and tensor parallelism for splitting a model across cards.
How a request flows
Putting the pieces together, one request travels like this. The frontend accepts an OpenAI-style call and hands it to the router. The router checks which workers already hold matching KV blocks and how loaded each one is, and picks a prefill worker. That worker processes the prompt and builds the KV cache. NIXL moves the cache to a decode worker, which streams tokens back through the frontend. Meanwhile the Planner watches pool utilization and latency and adds or removes workers in each pool. If a worker fails, the README notes request migration as part of its fault tolerance. Each step maps to a lever in our inference engineering overview.
When to pick Dynamo
Pick it when:
- Traffic is high enough that prefill and decode interfere, typically long prompts, agents or reasoning models.
- Many requests share prefixes (system prompts, RAG context, multi-turn chat), so KV-aware routing and offload pay back.
- You run many GPUs and want autoscaling driven by latency targets rather than raw utilization.
Skip it when you serve one model on one to eight GPUs with modest traffic. A single vLLM or SGLang server will be easier to run, and you can adopt Dynamo later because it wraps those same engines.
Performance
These are NVIDIA's published claims from the Dynamo README, each attributed to the source the README links. They are the project's numbers on its chosen workloads, not independent measurements, and we have not reproduced them.
- 7x higher throughput per GPU for DeepSeek R1 on GB200 NVL72 with Dynamo versus B200 without, per SemiAnalysis InferenceX.
- 2x faster time to first token from KV-aware routing on Qwen3-Coder 480B, per the Dynamo README.
- 80% fewer SLA breaches at 5% lower TCO with Planner autoscaling, per a presentation at Alibaba's APSARA 2025.
- 7x faster model startup from ModelExpress weight streaming (DeepSeek-V3 on H200); the README gives no link in that table row.
Treat these as upper-bound illustrations. Your workload, model and cluster decide what you get.
Run it on a cloud GPU
Start with a single multi-GPU box to learn the frontend and a backend worker, then scale out. Pick the card that matches your model size.
FAQ
Is NVIDIA Dynamo a replacement for vLLM?
No. Dynamo runs vLLM, SGLang or TensorRT-LLM as its workers and adds routing, disaggregation, cache tiering and scaling on top.
What is NIXL?
The NVIDIA Inference Transfer Library: a point-to-point library with one API for moving data between GPU memory, host memory and storage, used to move KV cache between workers.
What is disaggregated serving?
Running the prompt-processing phase (prefill) and the token-generation phase (decode) on separate GPU pools so each can be sized and scheduled for its own bottleneck.
Does Dynamo need Kubernetes?
The quickstart runs in a single container. Larger deployments use Kubernetes, and the README lists Grove for topology-aware scheduling there. Check the docs for your version.
Dynamo vs llm-d?
Both orchestrate engines with KV-aware routing and disaggregation. Dynamo is NVIDIA's framework; llm-d is a Kubernetes-native CNCF sandbox project built around the Gateway API Inference Extension.
Sources
- Dynamo README (components, backends, quickstart, performance table): https://github.com/ai-dynamo/dynamo
- Dynamo documentation (current version 1.5.1): https://docs.nvidia.com/dynamo/latest/index.html
- NVIDIA technical blog, Introducing NVIDIA Dynamo: https://developer.nvidia.com/blog/introducing-nvidia-dynamo-a-low-latency-distributed-inference-framework-for-scaling-reasoning-ai-models/
- DistServe paper (7.4x requests, 12.6x tighter SLO): https://arxiv.org/abs/2401.09670
- SemiAnalysis InferenceX: https://inferencex.semianalysis.com/
- ModelExpress: https://github.com/ai-dynamo/modelexpress