How to run Nemotron 3 Nano with vLLM and SGLang

Reasoning parser flags, enable_thinking toggle, 1M context setting and trust-remote-code notes for running NVIDIA Nemotron 3 Nano.

This guide covers serving the Nemotron 3 family, using the Nemotron-3-Nano-30B-A3B card as the reference.

Architecture

The card describes a Mamba2-Transformer hybrid mixture-of-experts with 52 layers: 23 MoE layers, 23 Mamba-2 layers and 6 grouped-query attention layers.

Engine support

  • vLLM 0.12.0 or newer. The card says it needs a custom reasoning parser file (nano_v3_reasoning_parser.py) supplied with the model.
  • SGLang: launch with --reasoning-parser nano_v3.
  • TRT-LLM: --reasoning_parser nano-v3 with the autodeploy backend.
  • Transformers: the chat template is integrated from v5.3.0.
  • The card only points to a model browser for llama.cpp quantizations and says nothing about Ollama, so check those engines yourself.

Reasoning toggle

Reasoning is on by default. Pass enable_thinking=False to apply_chat_template() to turn the trace off. The card suggests temperature 1.0 and top_p 1.0 for reasoning, and temperature 0.6 with top_p 0.95 for tool calling.

Long context

The model supports up to 1M tokens, but the default Hugging Face config is 256K because of memory. For 1M in vLLM, set VLLM_ALLOW_LONG_MAX_MODEL_LEN=1. No YaRN or rope scaling flags are documented.

Pitfalls

  • trust_remote_code is required because of the custom architecture.
  • Without the matching reasoning parser, thinking text appears inline in the reply.

Launch commands by Nemotron 3 size

One model per size. VRAM is for the precision each is published in; the GPU count comes from the cheapest live fit, and you set it with --tensor-parallel-size (vLLM) or --tp (SGLang).

NVIDIA-Nemotron-3-Nano-4B-BF16 (4.0B)

Needs about 8.9 GB of VRAM at BF16. Cheapest live fit: RTX 4070 Super at $0.121/hr.

vLLM

vllm serve nvidia/NVIDIA-Nemotron-3-Nano-4B-BF16 --tensor-parallel-size 1

SGLang

sglang serve --model-path nvidia/NVIDIA-Nemotron-3-Nano-4B-BF16 --tp 1

Generic example, not from the model's own docs: adjust flags (quantization, context length, parallelism) for your setup.

NVIDIA-Nemotron-3-Nano-30B-A3B-BF16 (31.6B)

Needs about 70.6 GB of VRAM at BF16. Cheapest live fit: A100 at $1.21/hr.

vLLM

vllm serve nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16 --tensor-parallel-size 1

SGLang

sglang serve --model-path nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16 --tp 1

Generic example, not from the model's own docs: adjust flags (quantization, context length, parallelism) for your setup.

NVIDIA-Nemotron-3-Super-120B-A12B-BF16 (123.6B)

Needs about 276 GB of VRAM at BF16. Cheapest live fit: 6× RTX A6000 at $2.18/hr.

vLLM

vllm serve nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16 --tensor-parallel-size 6

SGLang

sglang serve --model-path nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16 --tp 6

Generic example, not from the model's own docs: adjust flags (quantization, context length, parallelism) for your setup.

NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 (560.5B)

Needs about 1253 GB of VRAM at BF16. No live GPU fit right now.

vLLM

vllm serve nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16

SGLang

sglang serve --model-path nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16

Generic example, not from the model's own docs: adjust flags (quantization, context length, parallelism) for your setup.

Launch on Aquanode

Aquanode sells GPU pods billed per second, not a hosted inference API. Rent a pod sized to the model above, open a terminal on it or save a command as a startup script, then connect to the endpoint it serves.

Sources

Updated 2026-10-07.

Related

Submit the job. Everything after that is ours.

Sign up in 60 seconds. Pay for the GPU minutes you actually use.

© 2026 Aquanode. All rights reserved.

All trademarks, logos and brand names are the property of their respective owners.