LLM Inference Engines Compared: vLLM, SGLang and More (2026)

Back
Team Aquanode

Team Aquanode

Sarthak Vaish

Updated OCTOBER 8, 2026Published OCTOBER 8, 2026

An LLM inference engine is the software that loads a model onto hardware and turns prompts into tokens: it schedules requests, manages the KV cache, runs the attention and matrix-multiply kernels, and usually exposes an HTTP API. The right engine depends on where you run (laptop or datacenter GPU), how many users you serve, and whether you are on NVIDIA hardware, and this guide maps all of them in one place.

Below is a single comparison table built from each project's own documentation, then a short section per engine with a link to our deep dive, a how-to-choose section, and an FAQ. Details change fast, so every row cites the version or page it came from in the sources list.

TL;DR

  • Serving many users on GPUs: vLLM or SGLang are the open-source defaults. Both are Apache-2.0, expose OpenAI-compatible APIs, and run on more than NVIDIA hardware.
  • Squeezing a fixed NVIDIA fleet: TensorRT-LLM is NVIDIA's own engine with FP8 and FP4 recipes, and it is the one to benchmark when you are committed to NVIDIA GPUs.
  • Running models on your own machine: Ollama, LM Studio and llama.cpp are the simple routes, and all three can use a CPU, Apple silicon or a consumer GPU.
  • Many GPUs and many nodes: NVIDIA Dynamo and llm-d sit above engines and handle routing, prefill/decode disaggregation and cache offload.
  • Enterprise packaging: NVIDIA NIM wraps optimized engines in a container with a stable API, under NVIDIA's licensing.

Three layers: kernels, engines, orchestrators

People use "inference engine" loosely, so it helps to separate three layers.

  1. Kernel libraries supply fast attention, GEMM and MoE operations. FlashInfer is a kernel library, not a server; its README lists SGLang, vLLM, TensorRT-LLM and others as integrators. The same is true of FlashAttention (see FlashAttention 2 vs 3).
  2. Engines load a model, batch requests, manage the KV cache and serve an API. vLLM, SGLang, TensorRT-LLM, llama.cpp and Ollama live here. Their core techniques, PagedAttention and continuous batching, are covered in their own post.
  3. Orchestrators coordinate many engine instances across GPUs and nodes. Dynamo and llm-d are here, and Triton Inference Server plays a related role as a general model server.

The metrics you will compare engines on are covered in TTFT and tokens per second, and the broader craft is in inference engineering. For the basics, start with what is AI inference.

Comparison table

"Best for" is our summary. Everything else is from the project's own README or documentation, as cited in Sources. "Not stated" means the page we read did not say, not that the feature is absent.

EngineBest forHardwareQuantization formatsOpenAI-compatible APILicense
vLLMHigh-throughput serving of open modelsNVIDIA, AMD, Intel GPUs, CPUs; plugins for TPUs, Gaudi, Ascend and moreFP8, MXFP8/MXFP4, NVFP4, INT8, INT4, GPTQ/AWQ, GGUF, compressed-tensors and moreYesApache-2.0
SGLangAgentic and structured workloads, large-scale servingNVIDIA (A100 to B300 and select RTX), AMD Instinct, TPU, Intel, Apple silicon, AscendGPTQ/AWQ (and Marlin), FP8, INT8, FP4/MX formats, GGUF, bitsandbytes and moreYes ("Hugging Face and OpenAI APIs")Apache-2.0
TensorRT-LLMMaximum efficiency on NVIDIA GPUsNVIDIA only (Blackwell, Hopper, Ada, Ampere per its matrix)FP4/NVFP4, FP8 variants, INT4 AWQ, GPTQ W4A16/W4A8, FP8 and NVFP4 KV cacheYes (trtllm-serve)Apache-2.0
llama.cppLocal and edge inference, CPU and mixed CPU-GPUCPUs (x86, ARM, RISC-V), Apple silicon, CUDA, HIP, Vulkan, SYCL and more1.5-bit to 8-bit integer quantization (GGUF models)Yes (llama-server)MIT
OllamaOne-command local models and app integrationNVIDIA (compute capability 5.0+), AMD ROCm, Apple Metal, VulkanImports Safetensors and GGUF; does not quantize GGUF on importYes, a subsetMIT
LM StudioDesktop app with a local servermacOS, Windows, Linux (x64 and ARM64 on Windows)GGUF (llama.cpp) and MLX (Apple silicon)Yes ("OpenAI-like endpoints")Proprietary (free-to-use terms not stated in the terms page)
NVIDIA NIMPackaged, supported containers on NVIDIA GPUsNVIDIA GPUs, compute capability above 7.0Per model container; see NVIDIA's supported-models pageYes (works with the OpenAI client)NVIDIA Developer Program or AI Enterprise license
Triton Inference ServerServing many model types behind one serverNVIDIA GPUs, x86 and ARM CPUs, AWS InferentiaDepends on the backendNot stated in its READMEBSD-3-Clause
NVIDIA DynamoDatacenter-scale multi-node servingWhatever its backend engines supportInherited from vLLM, SGLang or TensorRT-LLMYes (frontend)Apache-2.0
llm-dKubernetes-native distributed inferenceTargets most accelerators; benchmarks on H100, H200, B200, MI300X, TPU, Intel XPUInherited from the model server (vLLM, SGLang)Not stated in its READMEApache-2.0

Notes on the table:

  • vLLM's hardware and quantization cells come from its README. Its documentation adds a hardware support matrix; for example, AWQ is marked supported on Turing through Hopper but not on Volta or AMD GPUs, while FP8 W8A8 via LLM Compressor is marked supported on Ada, Hopper and AMD GPUs. Check that matrix for your card before choosing a format.
  • TensorRT-LLM's matrix marks NVFP4 and MXFP4 as supported on Blackwell and not on Hopper, Ada or Ampere, and FP8 per-tensor as supported on Blackwell, Rubin, Hopper and Ada.
  • The two NVIDIA-only rows (TensorRT-LLM, NIM) are the ones where "hardware" means a hard requirement. The others run on several vendors.

The engines, one by one

vLLM

vLLM's README describes high serving throughput, PagedAttention for KV-cache memory, continuous batching, chunked prefill, prefix caching, structured outputs with xgrammar or guidance, and tensor, pipeline, data, expert and context parallelism. It serves an OpenAI-compatible API (and, per the README, an Anthropic Messages API and gRPC). The documented install and serve commands:

uv venv --python 3.12 --seed
source .venv/bin/activate
uv pip install vllm --torch-backend=auto
vllm serve Qwen/Qwen2.5-1.5B-Instruct

The server listens on port 8000 by default. Read our long-form vLLM guide for how it works, then vLLM Docker install for a production container, and vLLM vs Ollama if you are deciding between a server and a desktop tool.

SGLang

SGLang is an Apache-2.0 engine its README describes as optimized for agentic workloads, RL rollouts and large-scale serving. It supports text, vision-language and diffusion models, hierarchical KV caching (HiCache) and speculative decoding tooling. Its docs list RadixAttention, prefix caching and multi-GPU parallelism, and its quantization page lists offline and online methods. Documented commands:

uv pip install --prerelease=allow sglang
sglang serve MODEL_PATH --host 0.0.0.0 --port 30000

The SGLang quantization page recommends offline quantization over online quantization for performance and convenience. See the SGLang guide, and the existing vLLM vs TensorRT-LLM vs SGLang comparison.

TensorRT-LLM

TensorRT-LLM is NVIDIA's open-source library for efficient LLM inference on NVIDIA GPUs. Its README lists custom kernels, prefill-decode disaggregation, wide expert parallelism, speculative decoding, a PyTorch-based Python LLM API, and integration with Dynamo and Triton. trtllm-serve starts a server its docs call OpenAI compatible, serving /v1/models, /v1/completions and /v1/chat/completions:

trtllm-serve <model> [--tp_size <tp> --pp_size <pp> --ep_size <ep> --host <host> --port <port>]

Its quantization page says the default PyTorch backend supports FP4 and FP8 on the latest Blackwell and Hopper GPUs. Details are in the TensorRT-LLM guide.

llama.cpp

llama.cpp is a C/C++ engine that treats Apple silicon as a first-class target and also supports x86 CPUs (AVX, AVX2, AVX512, AMX), CUDA, HIP, Vulkan, SYCL and other backends. It offers 1.5-bit to 8-bit integer quantization and an OpenAI-compatible server. From its README:

llama serve -hf ggml-org/Qwen3.5-0.8B-GGUF

That is the right tool for CPU-only boxes, mixed CPU/GPU offload and edge devices. See the llama.cpp guide. For file formats, see the GGUF glossary entry.

Ollama

Ollama wraps llama.cpp (its README lists it under supported backends) in a one-line install and a model library. Its docs list NVIDIA GPUs with compute capability 5.0 and driver 550 or newer, AMD ROCm v7 on Linux, Apple Metal, and Vulkan on Windows and Linux. It exposes a subset of the OpenAI API at http://localhost:11434/v1/, including chat completions, completions, embeddings and a responses endpoint.

curl -fsSL https://ollama.com/install.sh | sh
ollama run gemma4

See what is Ollama, Ollama vs LM Studio and vLLM vs Ollama.

LM Studio

LM Studio is a desktop app for running open models locally. Its docs say it runs on macOS, Windows and Linux using llama.cpp, adds MLX on Apple silicon, serves OpenAI-like endpoints locally and on the network, and ships a headless version (llmster) and an lms CLI. It is proprietary software under its own terms, so check them if you plan commercial use. See Ollama vs LM Studio.

NVIDIA NIM

NIM for LLMs is a container that packages an optimized model server. NVIDIA's getting-started page shows an NGC-based docker run with --runtime=nvidia --gpus all, port 8000 and a mounted cache, and calls it from the OpenAI Python client by changing the base_url. Prerequisites listed there include an NVIDIA GPU with compute capability above 7.0 (8.0 for bfloat16), driver release 580 or later and Docker 23.0.1 or newer. Downloading and deploying the container requires the free NVIDIA Developer Program or an NVIDIA AI Enterprise license. Read the NIM guide.

Triton Inference Server

Triton is a general model server with HTTP/REST and gRPC protocols, dynamic batching, concurrent model execution, model ensembles and multiple backends (TensorRT, PyTorch, ONNX, OpenVINO, Python and more). It is BSD-3-Clause licensed and part of NVIDIA AI Enterprise. For LLMs it is most often paired with TensorRT-LLM, which TensorRT-LLM's README lists as an integration. See the Triton guide.

NVIDIA Dynamo

Dynamo describes itself as an open-source, datacenter-scale inference stack that sits above engines. It supports vLLM, SGLang and TensorRT-LLM backends and ships an OpenAI-compatible frontend, disaggregated prefill and decode, KV-aware routing, a KV block manager that offloads cache across GPU, CPU, SSD and remote storage, and an SLA-driven planner. Its README quickstart runs the frontend and a vLLM worker inside a container. See the Dynamo guide.

llm-d

llm-d is a CNCF Sandbox project and a distributed inference serving stack for Kubernetes. It adds prefix-cache-aware and load-aware routing, prefill/decode disaggregation with wide expert parallelism, tiered KV-cache offloading and SLO-aware autoscaling on top of model servers, naming vLLM and SGLang. See the llm-d guide.

How to choose

Start from your constraint, not from the benchmark chart.

If you...Start withWhy
Want a model running on your laptop in a minuteOllama or LM StudioInstallers, model libraries and a local OpenAI-style endpoint
Need CPU, Apple silicon or mixed CPU/GPU inferencellama.cppBroadest backend list, low-bit quantization
Serve an open model to many concurrent users on GPUsvLLM or SGLangContinuous batching, prefix caching, parallelism, OpenAI APIs
Run only NVIDIA GPUs and want the most from FP8 or FP4TensorRT-LLMNVIDIA's own recipes for Hopper and Blackwell
Need a supported container and stable APINIMNVIDIA packages and licenses it
Run a multi-node fleet with long promptsDynamo or llm-dRouting and disaggregated prefill/decode
Serve LLMs next to classic ML modelsTritonMany backends, one server

Then check four things before you commit:

  1. Model support. Look up your exact model in each project's supported-models list before you benchmark anything.
  2. Quantization on your GPU. Matrix support differs by architecture, as the notes under the table show. FP8, FP4, AWQ and GPTQ are not equally fast everywhere. See NVFP4 vs MXFP4.
  3. Memory. Weights plus KV cache decide what fits. Use how much VRAM do I need for LLMs and the KV cache entry.
  4. Your own traffic. Engines publish benchmarks on their own workloads. Measure your prompt and output lengths and concurrency. We do not quote cross-engine speed numbers here because the projects publish them under different conditions, and none of them are comparable out of the box.

For the GPU side of the decision, see best GPU for LLM inference, best GPU for AI and the H100, H200, L40S and RTX 5090 pages.

Run it on a cloud GPU

Pick an engine from the table and test it on your own model and prompts. The live box below shows what is available to rent right now.

FAQ

What is the best LLM inference engine?

There is no single best one. For multi-user GPU serving the common open-source choices are vLLM and SGLang, for NVIDIA-only fleets TensorRT-LLM is worth benchmarking, and for local use Ollama, LM Studio or llama.cpp are simpler.

Is Ollama an inference engine?

It is a local runtime and model manager. Its README lists llama.cpp as its backend, so the engine underneath is llama.cpp. It adds installers, a model library and an API, including a subset of the OpenAI API.

Which engines expose an OpenAI-compatible API?

vLLM, SGLang, TensorRT-LLM (trtllm-serve), llama.cpp (llama-server), Ollama (a subset), LM Studio and Dynamo (frontend) document one. NIM is called through the OpenAI client. Triton's and llm-d's READMEs did not mention one in the pages we read.

Do I need Dynamo or llm-d?

Only when one engine on one node is no longer enough: many replicas, long shared prefixes, or splitting prefill and decode across GPU pools. They orchestrate engines such as vLLM, SGLang and (for Dynamo) TensorRT-LLM rather than replace them.

Is FlashInfer an inference engine?

No. Its README calls it a library and kernel generator for inference, and engines such as SGLang, vLLM and TensorRT-LLM call it for attention, GEMM and MoE kernels.

Which engine works on AMD GPUs?

vLLM lists AMD GPUs, SGLang lists AMD Instinct MI300X to MI355X, llama.cpp supports HIP, and Ollama documents ROCm. TensorRT-LLM and NIM are NVIDIA only.

Sources

#llm inference#inference engines#llm serving#vllm#sglang#tensorrt-llm#llama.cpp

Submit the job. Everything after that is ours.

Sign up in 60 seconds. Pay for the GPU minutes you actually use.

© 2026 Aquanode. All rights reserved.

All trademarks, logos and brand names are the property of their respective owners.