Triton Inference Server: Setup, Batching, LLMs (2026)

Back
Team Aquanode

Team Aquanode

Sarthak Vaish

Updated OCTOBER 8, 2026Published OCTOBER 8, 2026

NVIDIA Triton Inference Server is open-source serving software that loads models from many frameworks (TensorRT, PyTorch, ONNX, OpenVINO, Python and more) from a model repository and serves them over HTTP and gRPC, batching requests on the fly. It is the general-purpose layer of the NVIDIA stack: you reach for it when you serve several kinds of model from one server, and for LLMs it usually sits in front of a dedicated engine such as TensorRT-LLM or vLLM.

This guide covers what Triton does, how to start it with Docker, how dynamic batching is configured, and how it relates to the newer LLM-specific stacks. It is part of our guide to LLM inference engines.

TL;DR

  • Triton is a model server, not an LLM engine. It schedules and batches requests; a backend (TensorRT, ONNX Runtime, PyTorch, Python, TensorRT-LLM, vLLM) runs the model.
  • Pick it when you serve a mix of models (vision, speech, embeddings, rankers, LLMs) behind one API with shared metrics and batching.
  • For a single LLM, a dedicated engine like vLLM or SGLang is simpler. For large-scale LLM serving with disaggregated prefill and decode, look at NVIDIA Dynamo.
  • The current release at the time of writing is Triton 2.73.0 (NGC container 26.09).

What Triton Inference Server is

According to the project README, Triton is "open source inference serving software" that deploys models from many frameworks and runs them across cloud, data center, edge and embedded targets. It runs on NVIDIA GPUs, x86 and ARM CPUs and AWS Inferentia, though not every backend runs on every platform, so check the project's backend support matrix before you plan a deployment. It is also part of NVIDIA AI Enterprise.

The idea is separation of concerns. The model author exports a model in whatever format suits them. The server owner points Triton at a directory of models. Clients talk to one protocol, whether the model behind it is a ResNet, a BERT ranker or a chat model.

How it works

The model repository

Triton serves whatever it finds in a model repository: a directory with one subfolder per model, each holding numbered version folders and an optional config.pbtxt. The server loads, unloads and updates models from that directory. The quickstart even runs with --model-control-mode explicit, which lets you load and unload named models at runtime through the API instead of loading everything at startup.

Backends

A backend is the code that executes a model. Triton's README lists TensorRT, PyTorch, ONNX, OpenVINO and Python among others (including RAPIDS FIL for tree models). The Python backend is the escape hatch: any Python code can be a model, which is how people wrap pre- and post-processing or non-standard models.

Dynamic batching

GPUs are far more efficient on batches than on single requests, but real clients send one request at a time. Triton's dynamic batcher, in the project's words, "allows inference requests to be combined by the server" so batches are formed on the fly for stateless models. You enable it per model in config.pbtxt. The documented snippets are:

dynamic_batching { }
dynamic_batching {
  preferred_batch_size: [ 4, 8 ]
}
dynamic_batching {
  max_queue_delay_microseconds: 100
}

Two pieces of guidance come straight from the batcher docs. First, a preferred batch size "should only be configured if that batch size results in significantly higher performance", so most models should leave it unset. Second, for the delay, "try increasing delay values until the latency budget is exceeded to see the impact on throughput." The batcher holds requests only as long as no request is delayed longer than the configured value, then sends whatever it has.

A minimal TensorRT model entry in the model configuration docs looks like this (the docs also note that a model with max_batch_size greater than 1 and no scheduler section gets the dynamic batcher by default for TensorRT models):

platform: "tensorrt_plan"
max_batch_size: 8

Ensembles and business logic scripting

An ensemble chains several models into a pipeline served as one model, for example tokenizer, model, detokenizer. Business Logic Scripting (BLS) does a similar job from Python code when the pipeline needs loops or conditions.

Metrics

Triton exposes Prometheus-style metrics covering GPU utilization, throughput and latency. In the default Docker quickstart the container publishes ports 8000, 8001 and 8002: HTTP, gRPC and metrics respectively.

Quickstart with Docker

These commands come from Triton's own quickstart and README. You need a machine with Docker and the NVIDIA Container Toolkit. See the vLLM Docker guide for the host setup that applies to any GPU container.

Fetch the example models:

cd docs/examples
./fetch_models.sh

Launch the server from the NGC container. The quickstart uses a placeholder for the tag; the README's current release is 26.09:

docker run --gpus=1 --rm -p8000:8000 -p8001:8001 -p8002:8002 \
  -v/full/path/to/docs/examples/model_repository:/models \
  nvcr.io/nvidia/tritonserver:<xx.yy>-py3 \
  tritonserver --model-repository=/models

Replace the tag placeholder with the release you want, for example 26.09. Check readiness:

curl -v localhost:8000/v2/health/ready

A 200 means the server is ready. To send a test request, pull the client SDK image and run its example client:

docker pull nvcr.io/nvidia/tritonserver:<xx.yy>-py3-sdk
docker run -it --rm --net=host nvcr.io/nvidia/tritonserver:<xx.yy>-py3-sdk
/workspace/install/bin/image_client -m densenet_onnx -c 3 -s INCEPTION /workspace/images/mug.jpg

Using Triton for LLMs

Triton does not generate text by itself. For LLMs you pair it with an LLM backend, most often the TensorRT-LLM backend. The tensorrtllm_backend repository describes its goal as letting you "serve TensorRT-LLM models with Triton Inference Server", and notes that the Triton backend source and tests now live in the TensorRT-LLM repository under the triton_backend directory. Its getting-started section points to the PyTorch-backend LLM API route, which it describes as serving "any HuggingFace model directly, no engine compilation required."

Two practical notes from the Triton release notes: the 2.73.0 and 2.72.0 release notes both state that the Triton TensorRT-LLM Backend container image is not included in those releases, so check the TensorRT-LLM project for its own container before you plan around that combination. And the 26.09 containers add initial support for the NVIDIA Rubin architecture.

Why would you still run an LLM through Triton rather than directly through an engine?

  • You already run Triton for other models and want one gateway, one metrics stack and one deployment pattern.
  • You need ensembles that combine an LLM with an embedding model, a classifier or a reranker.
  • You want NVIDIA AI Enterprise support for the serving layer.

Why might you not?

  • Continuous batching, paged KV cache and prefix caching live in the LLM engine, not in Triton's generic dynamic batcher. See paged attention and continuous batching for why that matters.
  • A single-model deployment gains little from the extra layer. vLLM and TensorRT-LLM each expose an OpenAI-compatible server on their own.

GPU and VRAM requirements

Triton itself needs almost no GPU memory. What you need is set by the models you load. As a computed rule, weights take parameters times bytes per parameter: a 7B model is about 14 GB at FP16, 7 GB at INT8 and 3.5 GB at INT4, before KV cache and runtime overhead. Use the VRAM calculator and our VRAM sizing guide for the full picture. Small vision and embedding models fit on an L40S; a mixed fleet with one or two 7-13B LLMs is comfortable on an H100.

Triton vs the alternatives

NeedBetter fit
Many model types behind one APITriton
One LLM, fastest path to an endpointvLLM or SGLang
Maximum NVIDIA-specific LLM performanceTensorRT-LLM
Packaged, supported NVIDIA containersNVIDIA NIM
Multi-node LLM serving with disaggregated prefill and decodeNVIDIA Dynamo

Our vLLM vs TensorRT-LLM vs SGLang comparison covers the engines underneath.

Performance

We are not quoting a Triton benchmark. NVIDIA's project documentation recommends tuning dynamic batching by raising the queue delay until you hit your latency budget, and measuring with its own tooling on your own model. Throughput depends on the model, the backend, the batch settings and the hardware, so measure on the GPU you plan to run. For the metrics to track, see TTFT and tokens per second.

Run it on a cloud GPU

Triton runs in a standard NVIDIA container, so any GPU box with Docker and the NVIDIA Container Toolkit works. Pick a card sized to the models in your repository.

FAQ

Is Triton Inference Server free?

The software is open source and free to use. NVIDIA also offers it as part of NVIDIA AI Enterprise, which adds enterprise support.

Does Triton support LLMs?

Yes, through LLM backends such as TensorRT-LLM and vLLM. Triton provides the server, scheduling and protocols; the backend provides token generation.

What ports does Triton use?

The quickstart publishes 8000 (HTTP), 8001 (gRPC) and 8002 (metrics). The readiness check is localhost:8000/v2/health/ready.

Triton vs vLLM: which should I use?

If you serve one LLM, vLLM is the shorter path. If you serve many model types behind one gateway, or you need ensembles, Triton is the better fit, and it can run an LLM backend inside it.

What does dynamic batching do?

It combines individual requests into batches on the server so the GPU works on several at once. You tune it with a maximum queue delay and, rarely, preferred batch sizes.

What replaced Triton for large LLM deployments?

Nothing replaced it for general serving. For large multi-node LLM serving, NVIDIA's newer framework is Dynamo, which sits above engines like vLLM, SGLang and TensorRT-LLM.

Sources

#llm inference#inference engines#triton#nvidia#dynamic batching#model serving

Submit the job. Everything after that is ours.

Sign up in 60 seconds. Pay for the GPU minutes you actually use.

© 2026 Aquanode. All rights reserved.

All trademarks, logos and brand names are the property of their respective owners.