NVIDIA NIM Guide: What It Is and How to Run It (2026)

Back
Team Aquanode

Team Aquanode

Sarthak Vaish

Updated OCTOBER 8, 2026Published OCTOBER 8, 2026

NVIDIA NIM (NVIDIA Inference Microservices) is a set of prebuilt, GPU-optimized containers that serve AI models behind an OpenAI-compatible API. You pull one image from NVIDIA's registry, run it with docker run --gpus=all, and get an endpoint on port 8000 with the engine and model profile already chosen for your GPU.

This guide covers what is inside a NIM, the exact run commands from NVIDIA's NIM for LLMs quickstart, the licensing rules for development versus production, and how NIM compares with running an engine like vLLM yourself. It is part of our guide to LLM inference engines.

TL;DR

  • NIM is a packaging layer: container image, model, tuned runtime and an OpenAI-compatible HTTP API in one unit. It is not a new inference engine.
  • Running one is a single docker run --gpus=all with a cache mount and port 8000. In NIM LLM 2.0.13, an NGC API key and docker login are optional for eligible public-catalog NIMs.
  • Licensing: NVIDIA says Developer Program members get free access to downloadable NIMs for development and testing on up to two nodes or 16 GPUs, and production needs an NVIDIA AI Enterprise license (with a free 90-day trial offered).
  • Pick NIM for the fastest path to a vendor-tested endpoint on NVIDIA GPUs with support. Pick vLLM, SGLang or TensorRT-LLM directly when you want full control over flags, versions and cost, or an open license for production.

What NVIDIA NIM is

A NIM is a container that bundles a model with an optimized inference runtime and a standard API. The NIM for LLMs documentation describes a quickstart that runs an LLM or vision-language model container and sends a first inference request. On start-up the container inspects your hardware and picks a model profile automatically, based on things like GPU count and architecture. You can override the choice with the NIM_MODEL_PROFILE environment variable. The docs' support matrix lists the correct image and version for each model and GPU.

NVIDIA's own engines sit underneath this packaging: TensorRT-LLM is the best-known one, and NIM is how NVIDIA delivers such an engine as a ready-to-run unit. The larger picture of how a NIM fits next to Triton and Dynamo is in our Triton Inference Server guide and NVIDIA Dynamo guide. We have not confirmed from the pages we read exactly which engine each NIM image uses internally, and NVIDIA picks the backend per model and profile, so check the model's page in the support matrix.

What you get compared with building your own container:

  • A tested combination. Model weights, engine version, quantization profile and driver requirements are validated together by NVIDIA.
  • One API. The server answers /v1/chat/completions and /v1/models, so OpenAI-style clients work by changing the base URL.
  • A cache directory. Downloaded model artifacts go in /opt/nim/.cache, which you mount from the host so restarts do not re-download.

What you give up: control over exact engine version and flags, and, for production, the cost of an enterprise license.

Prerequisites

The NIM docs keep prerequisites on a separate page (driver, Docker and the NVIDIA container toolkit); we did not read the contents of that page, so check it for your release. The basic requirements are the same as any GPU container: a Linux host, a recent NVIDIA driver, Docker with the NVIDIA container toolkit, and enough GPU memory for the model profile. Our VRAM guide and VRAM calculator help with the sizing.

Run a NIM

These steps follow the NIM LLM 2.0.13 quickstart and installation pages.

1. Log in to the registry (when you need to)

The installation page says that if you created an NGC API key, you authenticate to the NVIDIA Container Registry like this:

echo "$NGC_API_KEY" | docker login nvcr.io --username '$oauthtoken' --password-stdin

The literal username $oauthtoken means you authenticate with an API key instead of a username and password. Per the docs, login is optional for most public NIM images. You need it for Production Branch NIMs (their names carry a -pb plus version suffix), NIMs released before NIM LLM and VLM version 2.0.10, and private or gated images or model artifacts.

2. Set a cache directory and start the container

The quickstart uses the image nvcr.io/nim/meta/llama-3.1-8b-instruct:2.0.13, with a host cache directory mounted at /opt/nim/.cache and port 8000 published. Assembled from the quickstart's flags:

export NGC_API_KEY=<your-key>
export LOCAL_NIM_CACHE=~/.cache/nim
mkdir -p "$LOCAL_NIM_CACHE"

docker run --gpus=all \
  -e NGC_API_KEY=$NGC_API_KEY \
  -v "$LOCAL_NIM_CACHE:/opt/nim/.cache" \
  -p 8000:8000 \
  nvcr.io/nim/meta/llama-3.1-8b-instruct:2.0.13

Per the docs, you can omit the -e NGC_API_KEY line when the image and model support keyless access. NVIDIA's DGX Spark playbook also suggests adding --shm-size=16GB for that platform. Check the image tag against the support matrix for your model and version, since tags change with each release.

3. Test it

Find the served model name first:

curl http://localhost:8000/v1/models

Then send a chat completion, using the quickstart's model name and prompt:

curl http://localhost:8000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "meta/llama-3.1-8b-instruct",
    "messages": [{"role": "user", "content": "Hello! How are you?"}],
    "max_tokens": 100
  }'

Model-free NIM

The quickstart also documents a model-free image, nvcr.io/nim/nvidia/model-free-nim:2.0.13, that pulls weights from a source you choose. Its example sets NIM_MODEL_PATH, NIM_SERVED_MODEL_NAME (the example serves openai/gpt-oss-20b) and HF_TOKEN, with the same cache mount and port mapping. Model-source credentials such as HF_TOKEN can still be required even when no NGC key is.

Licensing: development versus production

This is where NIM differs most from open-source engines. From NVIDIA's developer blog post on NIM access:

  • Developer Program members get free access to downloadable NIM microservices for development and testing.
  • Usage under that program is capped at up to two nodes, or 16 GPUs.
  • For production, organizations are pointed to NVIDIA AI Enterprise, with a free 90-day license available when you are ready to use NIM in production.

NVIDIA's NIM FAQ draws the same line: Developer Program access is for prototyping, research, development and testing and does not include enterprise features or support, while a production deployment needs an NVIDIA AI Enterprise license. We did not find a license price on an NVIDIA page in our research and do not state one. Read the current terms before shipping anything user-facing. Individual models inside a NIM also carry their own model licenses, so check those separately.

NIM vs running vLLM, SGLang or TensorRT-LLM yourself

NVIDIA NIMSelf-managed engine
SetupPull one image, runChoose engine, version, flags, quantization
ControlVendor picks profile; override with NIM_MODEL_PROFILEFull control
HardwareNVIDIA GPUsvLLM and SGLang also support other vendors
Cost modelFree for development within limits; enterprise license for productionOpen-source licenses, you operate it
SupportNVIDIA support with the enterprise licenseCommunity

NIM is a good fit for teams that want a supported, validated endpoint and already buy NVIDIA software. A self-managed engine suits teams that need the newest model on day one, custom flags or non-NVIDIA hardware. Many teams prototype with NIM and move to a self-managed engine, or the reverse. See vLLM vs TensorRT-LLM vs SGLang for the engine-level comparison and Serving LLMs with vLLM for what a serving engine does internally.

GPU choice

NIM picks a profile for the GPU it finds, so the practical question is whether your model fits at all. An 8B model in FP16 has about 16 GB of weights (computed as 8 billion times 2 bytes), and a 70B model about 140 GB in FP16 or 70 GB at FP8, before KV cache. Cards such as the L40S handle small and mid-size models, while the H100 and H200 are the usual choices for 70B-class models. Whether a given NIM profile exists for your card is listed in NVIDIA's support matrix; we could not enumerate it from the pages we read.

Performance

NVIDIA publishes per-model performance tables in the NIM documentation, and we do not reproduce or estimate any here. Because a NIM wraps an engine, throughput depends on the profile the container selects, your precision, context length and concurrency. Measure on your own prompts and track time to first token and tokens per second; see TTFT and tokens per second.

Run it on a cloud GPU

All you need is a Linux box with a recent NVIDIA driver, Docker and the NVIDIA container toolkit, then the commands above. Live offers for the cards in this guide:

FAQ

What is NVIDIA NIM?

It is a set of prebuilt containers from NVIDIA that serve AI models with an optimized runtime behind an OpenAI-compatible API. It packages a tested model, engine and profile so you can start with one docker run command.

Is NVIDIA NIM free?

For development and testing, NVIDIA says Developer Program members get free access to downloadable NIMs on up to two nodes or 16 GPUs. Production use needs an NVIDIA AI Enterprise license, with a free 90-day trial offered.

Do I need an NGC API key?

Not always. In NIM LLM 2.0.13, the docs say the key and Docker login are optional for eligible public-catalog NIMs. You need one for Production Branch NIMs, for NIMs released before version 2.0.10, and for private or gated images.

What port does a NIM use?

Port 8000 in the quickstart, published with -p 8000:8000. The chat route is /v1/chat/completions and the model list is at /v1/models.

Is NIM faster than vLLM?

We cannot say. NVIDIA publishes its own performance tables, but we have not benchmarked NIM against vLLM, and any comparison depends on the model, precision and traffic. Test both on your workload.

Can I run a NIM on any GPU?

NIM targets NVIDIA GPUs and selects a model profile automatically for what it detects. The support matrix lists which model and GPU combinations have images.

Sources

#llm inference#inference engines#nvidia nim#nvidia#containers#ngc

Submit the job. Everything after that is ours.

Sign up in 60 seconds. Pay for the GPU minutes you actually use.

© 2026 Aquanode. All rights reserved.

All trademarks, logos and brand names are the property of their respective owners.