An LLM inference engine is the software that loads a model onto hardware and turns prompts into tokens: it schedules requests, manages the KV cache, runs the attention and matrix-multiply kernels, and usually exposes an HTTP API. The right engine depends on where you run (laptop or datacenter GPU), how many users you serve, and whether you are on NVIDIA hardware, and this guide maps all of them in one place.
Below is a single comparison table built from each project's own documentation, then a short section per engine with a link to our deep dive, a how-to-choose section, and an FAQ. Details change fast, so every row cites the version or page it came from in the sources list.
TL;DR
- Serving many users on GPUs: vLLM or SGLang are the open-source defaults. Both are Apache-2.0, expose OpenAI-compatible APIs, and run on more than NVIDIA hardware.
- Squeezing a fixed NVIDIA fleet: TensorRT-LLM is NVIDIA's own engine with FP8 and FP4 recipes, and it is the one to benchmark when you are committed to NVIDIA GPUs.
- Running models on your own machine: Ollama, LM Studio and llama.cpp are the simple routes, and all three can use a CPU, Apple silicon or a consumer GPU.
- Many GPUs and many nodes: NVIDIA Dynamo and llm-d sit above engines and handle routing, prefill/decode disaggregation and cache offload.
- Enterprise packaging: NVIDIA NIM wraps optimized engines in a container with a stable API, under NVIDIA's licensing.
Three layers: kernels, engines, orchestrators
People use "inference engine" loosely, so it helps to separate three layers.
- Kernel libraries supply fast attention, GEMM and MoE operations. FlashInfer is a kernel library, not a server; its README lists SGLang, vLLM, TensorRT-LLM and others as integrators. The same is true of FlashAttention (see FlashAttention 2 vs 3).
- Engines load a model, batch requests, manage the KV cache and serve an API. vLLM, SGLang, TensorRT-LLM, llama.cpp and Ollama live here. Their core techniques, PagedAttention and continuous batching, are covered in their own post.
- Orchestrators coordinate many engine instances across GPUs and nodes. Dynamo and llm-d are here, and Triton Inference Server plays a related role as a general model server.
The metrics you will compare engines on are covered in TTFT and tokens per second, and the broader craft is in inference engineering. For the basics, start with what is AI inference.
Comparison table
"Best for" is our summary. Everything else is from the project's own README or documentation, as cited in Sources. "Not stated" means the page we read did not say, not that the feature is absent.
| Engine | Best for | Hardware | Quantization formats | OpenAI-compatible API | License |
|---|---|---|---|---|---|
| vLLM | High-throughput serving of open models | NVIDIA, AMD, Intel GPUs, CPUs; plugins for TPUs, Gaudi, Ascend and more | FP8, MXFP8/MXFP4, NVFP4, INT8, INT4, GPTQ/AWQ, GGUF, compressed-tensors and more | Yes | Apache-2.0 |
| SGLang | Agentic and structured workloads, large-scale serving | NVIDIA (A100 to B300 and select RTX), AMD Instinct, TPU, Intel, Apple silicon, Ascend | GPTQ/AWQ (and Marlin), FP8, INT8, FP4/MX formats, GGUF, bitsandbytes and more | Yes ("Hugging Face and OpenAI APIs") | Apache-2.0 |
| TensorRT-LLM | Maximum efficiency on NVIDIA GPUs | NVIDIA only (Blackwell, Hopper, Ada, Ampere per its matrix) | FP4/NVFP4, FP8 variants, INT4 AWQ, GPTQ W4A16/W4A8, FP8 and NVFP4 KV cache | Yes (trtllm-serve) | Apache-2.0 |
| llama.cpp | Local and edge inference, CPU and mixed CPU-GPU | CPUs (x86, ARM, RISC-V), Apple silicon, CUDA, HIP, Vulkan, SYCL and more | 1.5-bit to 8-bit integer quantization (GGUF models) | Yes (llama-server) | MIT |
| Ollama | One-command local models and app integration | NVIDIA (compute capability 5.0+), AMD ROCm, Apple Metal, Vulkan | Imports Safetensors and GGUF; does not quantize GGUF on import | Yes, a subset | MIT |
| LM Studio | Desktop app with a local server | macOS, Windows, Linux (x64 and ARM64 on Windows) | GGUF (llama.cpp) and MLX (Apple silicon) | Yes ("OpenAI-like endpoints") | Proprietary (free-to-use terms not stated in the terms page) |
| NVIDIA NIM | Packaged, supported containers on NVIDIA GPUs | NVIDIA GPUs, compute capability above 7.0 | Per model container; see NVIDIA's supported-models page | Yes (works with the OpenAI client) | NVIDIA Developer Program or AI Enterprise license |
| Triton Inference Server | Serving many model types behind one server | NVIDIA GPUs, x86 and ARM CPUs, AWS Inferentia | Depends on the backend | Not stated in its README | BSD-3-Clause |
| NVIDIA Dynamo | Datacenter-scale multi-node serving | Whatever its backend engines support | Inherited from vLLM, SGLang or TensorRT-LLM | Yes (frontend) | Apache-2.0 |
| llm-d | Kubernetes-native distributed inference | Targets most accelerators; benchmarks on H100, H200, B200, MI300X, TPU, Intel XPU | Inherited from the model server (vLLM, SGLang) | Not stated in its README | Apache-2.0 |
Notes on the table:
- vLLM's hardware and quantization cells come from its README. Its documentation adds a hardware support matrix; for example, AWQ is marked supported on Turing through Hopper but not on Volta or AMD GPUs, while FP8 W8A8 via LLM Compressor is marked supported on Ada, Hopper and AMD GPUs. Check that matrix for your card before choosing a format.
- TensorRT-LLM's matrix marks NVFP4 and MXFP4 as supported on Blackwell and not on Hopper, Ada or Ampere, and FP8 per-tensor as supported on Blackwell, Rubin, Hopper and Ada.
- The two NVIDIA-only rows (TensorRT-LLM, NIM) are the ones where "hardware" means a hard requirement. The others run on several vendors.
The engines, one by one
vLLM
vLLM's README describes high serving throughput, PagedAttention for KV-cache memory, continuous batching, chunked prefill, prefix caching, structured outputs with xgrammar or guidance, and tensor, pipeline, data, expert and context parallelism. It serves an OpenAI-compatible API (and, per the README, an Anthropic Messages API and gRPC). The documented install and serve commands:
uv venv --python 3.12 --seed
source .venv/bin/activate
uv pip install vllm --torch-backend=auto
vllm serve Qwen/Qwen2.5-1.5B-Instruct
The server listens on port 8000 by default. Read our long-form vLLM guide for how it works, then vLLM Docker install for a production container, and vLLM vs Ollama if you are deciding between a server and a desktop tool.
SGLang
SGLang is an Apache-2.0 engine its README describes as optimized for agentic workloads, RL rollouts and large-scale serving. It supports text, vision-language and diffusion models, hierarchical KV caching (HiCache) and speculative decoding tooling. Its docs list RadixAttention, prefix caching and multi-GPU parallelism, and its quantization page lists offline and online methods. Documented commands:
uv pip install --prerelease=allow sglang
sglang serve MODEL_PATH --host 0.0.0.0 --port 30000
The SGLang quantization page recommends offline quantization over online quantization for performance and convenience. See the SGLang guide, and the existing vLLM vs TensorRT-LLM vs SGLang comparison.
TensorRT-LLM
TensorRT-LLM is NVIDIA's open-source library for efficient LLM inference on NVIDIA GPUs. Its README lists custom kernels, prefill-decode disaggregation, wide expert parallelism, speculative decoding, a PyTorch-based Python LLM API, and integration with Dynamo and Triton. trtllm-serve starts a server its docs call OpenAI compatible, serving /v1/models, /v1/completions and /v1/chat/completions:
trtllm-serve <model> [--tp_size <tp> --pp_size <pp> --ep_size <ep> --host <host> --port <port>]
Its quantization page says the default PyTorch backend supports FP4 and FP8 on the latest Blackwell and Hopper GPUs. Details are in the TensorRT-LLM guide.
llama.cpp
llama.cpp is a C/C++ engine that treats Apple silicon as a first-class target and also supports x86 CPUs (AVX, AVX2, AVX512, AMX), CUDA, HIP, Vulkan, SYCL and other backends. It offers 1.5-bit to 8-bit integer quantization and an OpenAI-compatible server. From its README:
llama serve -hf ggml-org/Qwen3.5-0.8B-GGUF
That is the right tool for CPU-only boxes, mixed CPU/GPU offload and edge devices. See the llama.cpp guide. For file formats, see the GGUF glossary entry.
Ollama
Ollama wraps llama.cpp (its README lists it under supported backends) in a one-line install and a model library. Its docs list NVIDIA GPUs with compute capability 5.0 and driver 550 or newer, AMD ROCm v7 on Linux, Apple Metal, and Vulkan on Windows and Linux. It exposes a subset of the OpenAI API at http://localhost:11434/v1/, including chat completions, completions, embeddings and a responses endpoint.
curl -fsSL https://ollama.com/install.sh | sh
ollama run gemma4
See what is Ollama, Ollama vs LM Studio and vLLM vs Ollama.
LM Studio
LM Studio is a desktop app for running open models locally. Its docs say it runs on macOS, Windows and Linux using llama.cpp, adds MLX on Apple silicon, serves OpenAI-like endpoints locally and on the network, and ships a headless version (llmster) and an lms CLI. It is proprietary software under its own terms, so check them if you plan commercial use. See Ollama vs LM Studio.
NVIDIA NIM
NIM for LLMs is a container that packages an optimized model server. NVIDIA's getting-started page shows an NGC-based docker run with --runtime=nvidia --gpus all, port 8000 and a mounted cache, and calls it from the OpenAI Python client by changing the base_url. Prerequisites listed there include an NVIDIA GPU with compute capability above 7.0 (8.0 for bfloat16), driver release 580 or later and Docker 23.0.1 or newer. Downloading and deploying the container requires the free NVIDIA Developer Program or an NVIDIA AI Enterprise license. Read the NIM guide.
Triton Inference Server
Triton is a general model server with HTTP/REST and gRPC protocols, dynamic batching, concurrent model execution, model ensembles and multiple backends (TensorRT, PyTorch, ONNX, OpenVINO, Python and more). It is BSD-3-Clause licensed and part of NVIDIA AI Enterprise. For LLMs it is most often paired with TensorRT-LLM, which TensorRT-LLM's README lists as an integration. See the Triton guide.
NVIDIA Dynamo
Dynamo describes itself as an open-source, datacenter-scale inference stack that sits above engines. It supports vLLM, SGLang and TensorRT-LLM backends and ships an OpenAI-compatible frontend, disaggregated prefill and decode, KV-aware routing, a KV block manager that offloads cache across GPU, CPU, SSD and remote storage, and an SLA-driven planner. Its README quickstart runs the frontend and a vLLM worker inside a container. See the Dynamo guide.
llm-d
llm-d is a CNCF Sandbox project and a distributed inference serving stack for Kubernetes. It adds prefix-cache-aware and load-aware routing, prefill/decode disaggregation with wide expert parallelism, tiered KV-cache offloading and SLO-aware autoscaling on top of model servers, naming vLLM and SGLang. See the llm-d guide.
How to choose
Start from your constraint, not from the benchmark chart.
| If you... | Start with | Why |
|---|---|---|
| Want a model running on your laptop in a minute | Ollama or LM Studio | Installers, model libraries and a local OpenAI-style endpoint |
| Need CPU, Apple silicon or mixed CPU/GPU inference | llama.cpp | Broadest backend list, low-bit quantization |
| Serve an open model to many concurrent users on GPUs | vLLM or SGLang | Continuous batching, prefix caching, parallelism, OpenAI APIs |
| Run only NVIDIA GPUs and want the most from FP8 or FP4 | TensorRT-LLM | NVIDIA's own recipes for Hopper and Blackwell |
| Need a supported container and stable API | NIM | NVIDIA packages and licenses it |
| Run a multi-node fleet with long prompts | Dynamo or llm-d | Routing and disaggregated prefill/decode |
| Serve LLMs next to classic ML models | Triton | Many backends, one server |
Then check four things before you commit:
- Model support. Look up your exact model in each project's supported-models list before you benchmark anything.
- Quantization on your GPU. Matrix support differs by architecture, as the notes under the table show. FP8, FP4, AWQ and GPTQ are not equally fast everywhere. See NVFP4 vs MXFP4.
- Memory. Weights plus KV cache decide what fits. Use how much VRAM do I need for LLMs and the KV cache entry.
- Your own traffic. Engines publish benchmarks on their own workloads. Measure your prompt and output lengths and concurrency. We do not quote cross-engine speed numbers here because the projects publish them under different conditions, and none of them are comparable out of the box.
For the GPU side of the decision, see best GPU for LLM inference, best GPU for AI and the H100, H200, L40S and RTX 5090 pages.
Run it on a cloud GPU
Pick an engine from the table and test it on your own model and prompts. The live box below shows what is available to rent right now.
FAQ
What is the best LLM inference engine?
There is no single best one. For multi-user GPU serving the common open-source choices are vLLM and SGLang, for NVIDIA-only fleets TensorRT-LLM is worth benchmarking, and for local use Ollama, LM Studio or llama.cpp are simpler.
Is Ollama an inference engine?
It is a local runtime and model manager. Its README lists llama.cpp as its backend, so the engine underneath is llama.cpp. It adds installers, a model library and an API, including a subset of the OpenAI API.
Which engines expose an OpenAI-compatible API?
vLLM, SGLang, TensorRT-LLM (trtllm-serve), llama.cpp (llama-server), Ollama (a subset), LM Studio and Dynamo (frontend) document one. NIM is called through the OpenAI client. Triton's and llm-d's READMEs did not mention one in the pages we read.
Do I need Dynamo or llm-d?
Only when one engine on one node is no longer enough: many replicas, long shared prefixes, or splitting prefill and decode across GPU pools. They orchestrate engines such as vLLM, SGLang and (for Dynamo) TensorRT-LLM rather than replace them.
Is FlashInfer an inference engine?
No. Its README calls it a library and kernel generator for inference, and engines such as SGLang, vLLM and TensorRT-LLM call it for attention, GEMM and MoE kernels.
Which engine works on AMD GPUs?
vLLM lists AMD GPUs, SGLang lists AMD Instinct MI300X to MI355X, llama.cpp supports HIP, and Ollama documents ROCm. TensorRT-LLM and NIM are NVIDIA only.
Sources
- vLLM README (license, hardware, quantization, features): https://github.com/vllm-project/vllm
- vLLM quantization docs (support matrix): https://docs.vllm.ai/en/latest/features/quantization/
- vLLM quickstart (install and serve commands): https://docs.vllm.ai/en/latest/getting_started/quickstart/
- SGLang README (license, hardware, features, install): https://github.com/sgl-project/sglang
- SGLang documentation (OpenAI and Hugging Face API compatibility, RadixAttention): https://docs.sglang.io/
- SGLang quantization docs: https://docs.sglang.io/advanced_features/quantization.html
- SGLang install page (serve command): https://docs.sglang.io/docs/get-started/install
- TensorRT-LLM README (license, hardware, features): https://github.com/NVIDIA/TensorRT-LLM
- TensorRT-LLM trtllm-serve docs: https://nvidia.github.io/TensorRT-LLM/commands/trtllm-serve/trtllm-serve.html
- TensorRT-LLM quantization docs (formats and hardware matrix): https://nvidia.github.io/TensorRT-LLM/features/quantization.html
- llama.cpp README (license, backends, quantization, server): https://github.com/ggml-org/llama.cpp
- Ollama README (license, install, backend): https://github.com/ollama/ollama
- Ollama OpenAI compatibility: https://docs.ollama.com/api/openai-compatibility
- Ollama GPU support: https://docs.ollama.com/gpu
- Ollama import docs: https://docs.ollama.com/import
- LM Studio docs: https://lmstudio.ai/docs/app
- LM Studio app terms: https://lmstudio.ai/app-terms
- NVIDIA NIM for LLMs getting started (version 1.15.0): https://docs.nvidia.com/nim/large-language-models/1.15.0/getting-started.html
- Triton Inference Server README: https://github.com/triton-inference-server/server
- NVIDIA Dynamo README: https://github.com/ai-dynamo/dynamo
- llm-d README: https://github.com/llm-d/llm-d
- FlashInfer README: https://github.com/flashinfer-ai/flashinfer