TensorRT-LLM Guide: Install, trtllm-serve, Hardware

Back
Team Aquanode

Team Aquanode

Sarthak Vaish

Updated OCTOBER 8, 2026Published OCTOBER 8, 2026

TensorRT-LLM is NVIDIA's open-source (Apache 2.0) library for optimizing LLM inference on NVIDIA GPUs, built on PyTorch with a Python LLM API and an OpenAI-compatible server called trtllm-serve. Install it with pip3 install tensorrt_llm on Ubuntu 24.04, or pull the NGC container nvcr.io/nvidia/tensorrt-llm/release, then run trtllm-serve "TinyLlama/TinyLlama-1.1B-Chat-v1.0".

This guide covers what the library does, the exact install and serve commands from its docs, supported hardware, and when to choose it over vLLM or SGLang. It is part of our guide to LLM inference engines.

TL;DR

  • TensorRT-LLM is NVIDIA-only. It supplies custom kernels for attention, GEMMs and MoE, plus runtime features such as prefill-decode disaggregation, wide expert parallelism and speculative decoding.
  • The modern workflow is PyTorch based: trtllm-serve MODEL or the Python LLM API, with no separate engine-build step in the quickstart.
  • Supported GPUs span Blackwell (B200, GB200, B300, GB300, DGX Spark), Hopper (H100, H200, GH200), Ada (L20, L40, L40S) and Ampere (A100), per the supported-hardware page.
  • Pick it when you run NVIDIA GPUs at scale and want vendor-tuned kernels and FP8 or FP4 quantized checkpoints. Pick vLLM or SGLang when you want portability across vendors or simpler operations.
  • It is under rapid development: the GitHub releases page currently shows release candidates (v1.3.0rc29, 29 September), and the README mentions 1.4.0rc0. Pin versions.

What TensorRT-LLM is

The README describes TensorRT-LLM as an open-source library for optimizing LLM and visual generation inference on NVIDIA GPUs. It is built on PyTorch and offers a Python LLM API. Its headline pieces, all listed in the README and docs index:

  • Custom kernels for common operations such as attention, GEMMs and mixture-of-experts layers.
  • Runtime optimizations: prefill-decode disaggregation, wide expert parallelism and speculative decoding.
  • Parallelism: the LLM API scales "from single-GPU to multi-GPU or multi-node deployments".
  • Ecosystem integration with NVIDIA Dynamo and Triton Inference Server (see our Dynamo guide and Triton Inference Server guide).
  • Modularity: pre-defined models can be customized in native PyTorch code.
  • Quantized models: FP8 and FP4 checkpoints are published on Hugging Face and linked from the README.
  • Visual generation: diffusion model support, listed as beta in the docs navigation.

The docs index also lists a KV cache system and KV cache compression, LoRA, sparse attention, guided decoding, multimodal support, an overlap scheduler, torch compile with prefill CUDA graphs, and Helix parallelism. Older tutorials describe a separate step that compiles a model into a TensorRT engine. The current quickstart does not require that, so follow the current docs rather than a 2024 blog post.

One more note from the README: anonymous usage telemetry is on by default and can be disabled with environment variables, a config file, the Python TelemetryConfig(disabled=True) option or the --no-telemetry CLI flag.

Supported hardware

The supported-hardware page says TensorRT-LLM supports the full spectrum of NVIDIA GPU architectures and lists:

ArchitectureGPUs named on the page
BlackwellB200, GB200, B300, GB300, DGX Spark
HopperH100, H200, GH200
Ada LovelaceL20, L40, L40S
AmpereA100

The README also mentions Jetson AGX Orin (initial support via JetPack 6.1, a November 2024 news item). For the cards most people rent, see the pages for the H100, H200, B200, L40S and A100. Consumer RTX cards are not listed on the supported-hardware page, so verify against the support matrix before planning around one.

Low-precision formats matter here. The quickstart warns to ensure your GPU supports FP8 before running the FP8 example. FP8 arrives with Hopper and Ada; FP4 is a Blackwell feature. Our glossary covers FP8 and FP4, and the blog explains NVFP4 vs MXFP4.

Install TensorRT-LLM

There are two routes in the docs: a pip install on a prepared machine, or the NGC container.

Option 1: pip

The Linux installation page lists these prerequisites:

  • OS: tested on Ubuntu 24.04.
  • Python: tested on Python 3.12.
  • CUDA: CUDA Toolkit 13.1, with the CUDA_HOME environment variable set. The cuda-compat-13-1 package may be needed depending on your driver.
  • PyTorch: version 2.10.0 with the CUDA 13.0 build.
  • System packages: OpenMPI headers, plus ZeroMQ if you use disaggregated serving.

Commands, copied from the page:

pip3 install torch==2.10.0 torchvision --index-url https://download.pytorch.org/whl/cu130
sudo apt-get -y install libopenmpi-dev
# Optional, only for disagg-serving:
sudo apt-get -y install libzmq3-dev
pip3 install --ignore-installed pip setuptools wheel && pip3 install tensorrt_llm

The page warns that on some systems (Ubuntu 22.04 is the example) pip may replace an existing CUDA 13.0 PyTorch with a CUDA 12.8 build, and gives a constraints-file workaround:

CURRENT_TORCH_VERSION=$(python3 -c "import torch; print(torch.__version__)")
echo "torch==$CURRENT_TORCH_VERSION" > /tmp/torch-constraint.txt
pip3 install --ignore-installed pip setuptools wheel && pip3 install tensorrt_llm -c /tmp/torch-constraint.txt

Because the PyPI wheel is built against one exact PyTorch version, the version coupling is tight. That is the main reason many teams prefer the container.

Option 2: the NGC container

The container images page gives the pull command, with the release tag in the version position (the page's example uses 1.3.0rc29):

docker pull nvcr.io/nvidia/tensorrt-llm/release:1.3.0rc29

The docs say to browse available tags on NGC, and to try a tag from a previous GitHub pre-release or release if yours is not available. Examples are installed in /app/tensorrt_llm/examples inside the release image. The documentation gives make -C docker release_run for a locally built release image but, in the pages we read, no ready-made docker run line for the NGC image. For running the NGC image, follow the container images page and the NGC catalog page for the release image rather than a command copied from a blog.

Serve a model with trtllm-serve

The quickstart serves a Hugging Face model with one command:

trtllm-serve "TinyLlama/TinyLlama-1.1B-Chat-v1.0"

For an FP8 checkpoint on a GPU that supports FP8:

trtllm-serve "nvidia/Qwen3-8B-FP8"

The server listens on port 8000 and speaks the OpenAI chat protocol. Test it with the quickstart's request:

curl -X POST http://localhost:8000/v1/chat/completions \
    -H "Content-Type: application/json" \
    -H "Accept: application/json" \
    -d '{
        "model": "TinyLlama/TinyLlama-1.1B-Chat-v1.0",
        "messages":[{"role": "system", "content": "You are a helpful assistant."},
                    {"role": "user", "content": "Where is New York? Tell me in a single sentence."}],
        "max_tokens": 32,
        "temperature": 0
    }'

The Python LLM API

For offline batch generation, the quickstart uses the LLM API:

from tensorrt_llm import LLM, SamplingParams

llm = LLM(model="TinyLlama/TinyLlama-1.1B-Chat-v1.0")
sampling_params = SamplingParams(temperature=0.8, top_p=0.95)

prompts = ["Hello, my name is", "The capital of France is"]
for output in llm.generate(prompts, sampling_params):
    print(
        f"Prompt: {output.prompt!r}, Generated text: {output.outputs[0].text!r}"
    )

The quickstart has a longer version with a main guard and a larger prompt list; the lines above are the core, with a short prompts list added for clarity. For tensor and pipeline parallel settings and the full list of trtllm-serve options, use the CLI reference in the docs for the release you installed, since flags change between release candidates.

GPU memory and sizing

The same rule applies as for any engine: weights plus KV cache plus overhead. Weight sizes are parameters times bytes per parameter (computed here):

Model sizeFP16 / BF16FP8FP4
8B16 GB8 GB4 GB
70B140 GB70 GB35 GB

Weights only; add KV cache for your context length and concurrency. Use the VRAM calculator for per-card numbers and see how much VRAM you need for LLMs. An FP8 70B model has 70 GB of weights, which is why the H200 with its larger memory is a common single-card target, while an H100 usually needs two cards for the same model.

TensorRT-LLM vs vLLM vs SGLang

TensorRT-LLMvLLMSGLang
VendorsNVIDIA onlyNVIDIA, AMD, Intel and othersNVIDIA, AMD, TPU, Intel, Apple and others
StrengthVendor-tuned kernels, FP8 and FP4 on new NVIDIA silicon, disaggregated servingBroad model and quantization coverage, simple opsPrefix reuse (RadixAttention), agent workloads
Start commandtrtllm-serve MODELvllm serve MODELsglang serve MODEL_PATH
LicenseApache 2.0Apache 2.0Apache 2.0

Choose TensorRT-LLM when you are committed to NVIDIA hardware, run the newest architectures (Blackwell FP4 especially) and can afford to track a fast-moving release stream. Choose vLLM or SGLang for faster time to a working endpoint or for non-NVIDIA hardware. Our deeper comparison is vLLM vs TensorRT-LLM vs SGLang, and the background on batching and KV memory is in Serving LLMs with vLLM.

Performance: what NVIDIA publishes

We do not publish our own TensorRT-LLM benchmarks. NVIDIA's README lists several vendor claims, which we repeat as theirs and without endorsing:

  • "Over 40,000 tokens per second on B200 GPUs" for Llama 4 (stated on the README without a link).
  • "24,000 tokens per second" for Meta Llama 3 (NVIDIA blog).
  • Blackwell breaking the 1,000 tokens-per-second-per-user barrier with Llama 4 Maverick (NVIDIA developer blog).
  • "Boost Llama 3.3 70B inference throughput 3x" with speculative decoding (NVIDIA developer blog).

Each depends on specific hardware, precision, sequence lengths and batch settings described in the linked posts. They are NVIDIA's numbers on NVIDIA's chosen workloads. To compare engines fairly, run the same prompts and concurrency on each one and record time to first token and tokens per second; see TTFT and tokens per second.

Run it on a cloud GPU

TensorRT-LLM needs a recent NVIDIA driver and, for the pip route, a matching CUDA toolkit; the container route needs only Docker with the NVIDIA container toolkit. Live offers for the cards in this guide:

FAQ

Is TensorRT-LLM free?

Yes. It is open source under Apache 2.0, per the README. The NVIDIA GPUs it runs on are the cost.

Does TensorRT-LLM run on AMD GPUs?

No. It targets NVIDIA GPUs only. The supported-hardware page lists Blackwell, Hopper, Ada Lovelace and Ampere parts.

How do I serve a model with TensorRT-LLM?

Run trtllm-serve "MODEL_ID", for example trtllm-serve "TinyLlama/TinyLlama-1.1B-Chat-v1.0". It exposes an OpenAI-compatible API on port 8000, as the quickstart shows.

Do I still need to build a TensorRT engine?

The current quickstart serves a Hugging Face model directly with trtllm-serve or the Python LLM API, with no separate engine build step in the commands. Older guides that describe an explicit build step predate the PyTorch workflow.

Which version should I use?

The releases page shows release candidates as the newest entries (v1.3.0rc29 on 29 September), and the docs are generated from a newer main branch. We could not confirm a stable 1.x tag from the pages we read, so pin the exact tag you tested.

Is TensorRT-LLM the same as NVIDIA NIM or Triton?

No. TensorRT-LLM is the engine library. NIM packages engines into prebuilt containers (see our NVIDIA NIM guide), and Triton is a serving framework that can host it (see the Triton guide).

Sources

#llm inference#inference engines#tensorrt-llm#nvidia#trtllm-serve#fp8

Submit the job. Everything after that is ours.

Sign up in 60 seconds. Pay for the GPU minutes you actually use.

© 2026 Aquanode. All rights reserved.

All trademarks, logos and brand names are the property of their respective owners.