SGLang is an open-source inference framework for large language, vision-language and diffusion models, built around reusing the KV cache across requests with a radix tree (RadixAttention). Install it with uv pip install --prerelease=allow sglang, start a server with sglang serve MODEL_PATH --host 0.0.0.0 --port 30000, and talk to it through an OpenAI-compatible API.
This guide covers what SGLang does differently, the commands straight from its docs, which hardware it supports, and how to decide between SGLang and vLLM. It is part of our guide to LLM inference engines.
TL;DR
- SGLang is a serving engine for LLMs, vision-language models and diffusion models, with an OpenAI-compatible HTTP API.
- Its signature idea is RadixAttention: the KV cache of finished requests is kept in a radix tree so later requests with a shared prefix skip recomputing it. That pays off most for agents, multi-turn chat and few-shot prompts.
- The current stable release at the time of writing is v0.5.21 (released 2 October, per the GitHub releases page).
- Pick SGLang when many requests share long prefixes or you run agent and RL rollout workloads. Pick vLLM when you want the broadest model and quantization coverage with the most conservative defaults. Both are good; test on your own traffic.
- Run it in the official
lmsysorg/sglangDocker image to avoid CUDA and PyTorch version conflicts.
What SGLang is
SGLang comes from the LMSYS group and the open-source community around it. The project README describes it as an inference framework for large language, vision-language and diffusion models, aimed at agentic workloads, RL rollouts and large-scale serving. It started as a pair of things: a frontend language for writing structured LLM programs (multiple calls, branching, constrained output) and a runtime that executes them efficiently. Today most people use the runtime as a plain server.
The pieces worth knowing, all listed in the project README:
- SGLang Diffusion: image and video generation built into the repo and the
sglangPython package. - HiCache: hierarchical KV caching across GPU memory, host memory and external storage.
- Speculative decoding: supported through SpecForge, which trains draft models for use with SGLang. See our glossary entry on speculative decoding.
- Audio: SGLang Omni serves text-to-speech and speech recognition models.
- Deployment tooling: routing and load balancing through projects such as SMG, Ray Serve and NVIDIA Dynamo (see the NVIDIA Dynamo guide).
How RadixAttention works
Every LLM request has a prefill phase that turns the prompt into keys and values stored in the KV cache. A naive server throws that cache away when the request ends. If the next request starts with the same system prompt, the same few-shot examples or the same chat history, the server recomputes identical work.
The LMSYS launch post describes the fix: instead of discarding the KV cache after a generation request, SGLang retains the cache for both prompts and generation results in a radix tree. A radix tree is a prefix tree with compressed edges, so it supports prefix search, insertion and eviction efficiently. Eviction follows an LRU policy that removes leaf nodes recursively, and a cache-aware scheduling policy tries to raise the hit rate. The frontend always sends full prompts; the runtime does the prefix matching and reuse on its own.
When does this matter?
- Agents and tool loops: each step re-sends the whole transcript plus one new message. The long prefix is shared.
- Multi-turn chat: the history is the shared prefix.
- Few-shot and RAG templates: the instructions and examples are identical across requests.
- Tree-of-thought and best-of-n sampling: many branches share a common trunk.
When requests are all unique short prompts, there is little to reuse and the benefit shrinks. Prefix caching exists in vLLM too (see our vLLM guide), and the broader memory-management ideas are covered in PagedAttention and continuous batching. The difference is emphasis: SGLang treats prefix reuse as the core of the design and ships a cache-aware scheduler around it. The v0.5.21 release notes say the prefix cache now runs on a Rust core by default, which shows how much of the project's effort goes into this component.
Supported hardware
The README lists these targets:
| Vendor | Supported hardware (per the SGLang README) |
|---|---|
| NVIDIA | A100, H100, H200, H800, H20, B200, B300, GB200, GB300, select RTX 30/40/50 series, RTX 6000 Ada and PRO 6000, DGX Spark, Jetson Orin |
| AMD | Instinct MI300X, MI325X, MI350X, MI355X |
| TPU v6e and v7 | |
| Intel | Arc and Arc Pro B-Series GPUs, Xeon CPUs |
| Apple | Macs via Metal and MLX |
| Huawei | Ascend A2, A3, 950PR/DT NPUs |
| Moore Threads | MTT S5000 |
For the NVIDIA datacenter cards that most people rent, see the pages for the H100, H200 and B200. The README does not publish a list of supported model families and points to the project Cookbook for model compatibility, so check the Cookbook for your exact model before you commit to a GPU.
Install SGLang
The install page offers pip or uv, nightly builds, source builds and Docker. These commands are copied from it.
With uv (recommended in the docs):
pip install --upgrade pip
pip install uv
uv venv --python 3.12
source .venv/bin/activate
uv pip install --prerelease=allow sglang
Nightly build:
uv pip install --prerelease=allow --index-strategy unsafe-best-match --extra-index-url https://docs.sglang.ai/whl/cu130/ sglang
From source (the docs pin the v0.5.21 tag):
git clone -b v0.5.21 https://github.com/sgl-project/sglang.git
cd sglang
pip install --upgrade pip
pip install -e "python"
Docker (the easiest route on a fresh GPU box):
docker pull lmsysorg/sglang:latest
HF_CACHE_DIR="/path/to/huggingface-cache"
docker run -it --gpus all \
--shm-size 32g \
-p 30000:30000 \
-v "$HF_CACHE_DIR:/root/.cache/huggingface" \
--ipc=host \
lmsysorg/sglang:latest bash
The container mounts your Hugging Face cache so weights survive restarts, publishes port 30000, and uses host IPC plus a 32 GB shared-memory segment for multi-GPU communication.
Launch a server and send a request
Inside the environment or container, the install page starts a server with sglang serve:
sglang serve MODEL_PATH --host 0.0.0.0 --port 30000
The older module form still appears in the docs (for example in the SkyPilot example):
python3 -m sglang.launch_server \
--model-path meta-llama/Llama-3.1-8B-Instruct \
--host 0.0.0.0 \
--port 30000
Replace the model with any supported Hugging Face repository. Gated models such as Llama need an HF_TOKEN in the environment.
The "Send a request" page shows two client styles. With curl, the server exposes the OpenAI-style chat route:
curl -s http://localhost:30000/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "meta-llama/Llama-3.1-8B-Instruct",
"messages": [{"role": "user", "content": "What is the capital of France?"}],
"max_tokens": 64
}'
With the OpenAI Python client you point the base URL at the server, as the docs do:
import openai
client = openai.Client(base_url="http://127.0.0.1:30000/v1", api_key="None")
response = client.chat.completions.create(
model="meta-llama/Llama-3.1-8B-Instruct",
messages=[{"role": "user", "content": "List 3 countries and their capitals."}],
temperature=0,
max_tokens=64,
)
print(response.choices[0].message.content)
The model name in the request must match the model the server loaded. The exact request shape above follows the docs' examples; the model name is swapped to match the launch command in this post.
GPU and VRAM requirements
SGLang does not publish a single VRAM requirement because it depends on the model, precision and context length. The rule is the same as for any engine: weights plus KV cache plus overhead. A back-of-envelope for the weights (computed as parameters times bytes per parameter):
| Model size | FP16 / BF16 (2 bytes) | FP8 / INT8 (1 byte) | INT4 (0.5 byte) |
|---|---|---|---|
| 8B | 16 GB | 8 GB | 4 GB |
| 70B | 140 GB | 70 GB | 35 GB |
These are weights only (computed), before the KV cache. A 70B model in FP8 is 70 GB of weights, so it needs more than one 80 GB card once you add cache, which is why people run it on two H100s or one H200 with its larger memory. For exact numbers per card use the VRAM calculator and our explainer on how much VRAM you need for LLMs. For multi-GPU serving SGLang supports tensor parallelism; check the docs of your release for the exact flag names, since they change between versions.
SGLang vs vLLM
Both engines serve an OpenAI-compatible API, support continuous batching, paged KV memory, quantization and multi-GPU parallelism. The differences are about emphasis:
| SGLang | vLLM | |
|---|---|---|
| Design center | Prefix reuse (RadixAttention), agent and RL workloads, structured programs | General-purpose serving, widest hardware and model coverage |
| Start command | sglang serve MODEL_PATH | vllm serve MODEL |
| Docker image | lmsysorg/sglang | vllm/vllm-openai |
| Extras | Built-in diffusion engine, HiCache, SpecForge | Large quantization menu, multi-LoRA, wide hardware support |
Choose SGLang when your traffic is full of shared prefixes, when you run multi-step agents, or when you want one engine for text and image generation. Choose vLLM when you need the longest list of supported quantization formats and hardware, or when your team already runs it. Many teams benchmark both on their own prompts; we cover the three-way tradeoff with TensorRT-LLM in vLLM vs TensorRT-LLM vs SGLang, and the vLLM side in Serving LLMs with vLLM and vLLM install and Docker. For a Triton-based stack see the Triton Inference Server guide.
Performance: what the project publishes
We do not run our own benchmarks in this post. The one headline number comes from the LMSYS launch post (January 2024): SGLang achieved "up to 5 times higher throughput" than baseline systems. The stated conditions: Llama-7B on one NVIDIA A10G (24 GB) in FP16; Mixtral-8x7B on eight A10G GPUs in FP16 with tensor parallelism of 8; baselines vLLM v0.2.5, Guidance v0.1.8 and Hugging Face TGI v1.3.0; nine workloads including MMLU, ReAct agents, Tree-of-Thought, JSON decoding and DSPy RAG. "Up to" is a peak, those baseline versions are old, and the result is for workloads that reward prefix reuse. Treat it as an explanation of why the design exists, not as a prediction for your hardware. The paper is on arXiv (2312.07104). The v0.5.21 notes also mention about 22 percent faster first-token latency on long prompts for DeepSeek-V4.1, which is a model-specific figure from the release notes.
To measure for yourself, send identical prompts to both engines at your real concurrency and compare time to first token and tokens per second; our TTFT and tokens per second post explains what to record.
Run it on a cloud GPU
Rent a box with a recent NVIDIA driver, install Docker with the NVIDIA container toolkit, and run the lmsysorg/sglang image above. The live offers for the cards used in this guide are below.
FAQ
Is SGLang better than vLLM?
Neither wins everywhere. SGLang is designed around prefix reuse and agent workloads, so it tends to shine when requests share long prefixes. vLLM has the broadest model, quantization and hardware coverage. Benchmark both on your own traffic.
What does RadixAttention do?
It keeps the KV cache of finished requests in a radix tree so a new request that starts with the same tokens can reuse the stored keys and values instead of running prefill again. It uses LRU eviction and a cache-aware scheduler, per the LMSYS launch post.
How do I start an SGLang server?
Run sglang serve MODEL_PATH --host 0.0.0.0 --port 30000, or the older python3 -m sglang.launch_server --model-path MODEL --host 0.0.0.0 --port 30000. The server exposes OpenAI-compatible routes under /v1.
Does SGLang run on AMD GPUs?
The README lists AMD Instinct MI300X, MI325X, MI350X and MI355X among supported hardware, along with TPUs, Intel GPUs, Apple Silicon and several NPUs.
Can SGLang generate images?
Yes. SGLang Diffusion is a built-in image and video generation engine, included in the repo and the sglang Python package, per the README.
Which GPU should I use?
Size the weights first (parameters times bytes per parameter), add KV cache for your context and concurrency, then pick the card. An 8B model fits comfortably on one 80 GB card in FP16; 70B models need FP8 or INT4 plus more than one GPU or a larger-memory card.
Sources
- SGLang README: https://github.com/sgl-project/sglang (raw: https://raw.githubusercontent.com/sgl-project/sglang/main/README.md)
- SGLang install docs: https://docs.sglang.io/get_started/install.html
- SGLang send-request docs: https://docs.sglang.io/basic_usage/send_request.html
- SGLang releases (v0.5.21): https://github.com/sgl-project/sglang/releases
- LMSYS launch post, RadixAttention and the 5x claim: https://www.lmsys.org/blog/2024-01-17-sglang/
- SGLang paper: https://arxiv.org/abs/2312.07104
- Weight-size table: computed here as parameters times bytes per parameter.