How to run Nemotron 3 Nano with vLLM and SGLang
Reasoning parser flags, enable_thinking toggle, 1M context setting and trust-remote-code notes for running NVIDIA Nemotron 3 Nano.
This guide covers serving the Nemotron 3 family, using the Nemotron-3-Nano-30B-A3B card as the reference.
Architecture
The card describes a Mamba2-Transformer hybrid mixture-of-experts with 52 layers: 23 MoE layers, 23 Mamba-2 layers and 6 grouped-query attention layers.
Engine support
- vLLM 0.12.0 or newer. The card says it needs a custom reasoning parser file (
nano_v3_reasoning_parser.py) supplied with the model. - SGLang: launch with
--reasoning-parser nano_v3. - TRT-LLM:
--reasoning_parser nano-v3with the autodeploy backend. - Transformers: the chat template is integrated from v5.3.0.
- The card only points to a model browser for llama.cpp quantizations and says nothing about Ollama, so check those engines yourself.
Reasoning toggle
Reasoning is on by default. Pass enable_thinking=False to apply_chat_template() to turn the trace off. The card suggests temperature 1.0 and top_p 1.0 for reasoning, and temperature 0.6 with top_p 0.95 for tool calling.
Long context
The model supports up to 1M tokens, but the default Hugging Face config is 256K because of memory. For 1M in vLLM, set VLLM_ALLOW_LONG_MAX_MODEL_LEN=1. No YaRN or rope scaling flags are documented.
Pitfalls
trust_remote_codeis required because of the custom architecture.- Without the matching reasoning parser, thinking text appears inline in the reply.
Launch commands by Nemotron 3 size
One model per size. VRAM is for the precision each is published in; the GPU count comes from the cheapest live fit, and you set it with --tensor-parallel-size (vLLM) or --tp (SGLang).
NVIDIA-Nemotron-3-Nano-4B-BF16 (4.0B)
Needs about 8.9 GB of VRAM at BF16. Cheapest live fit: RTX 4070 Super at $0.121/hr.
vLLM
vllm serve nvidia/NVIDIA-Nemotron-3-Nano-4B-BF16 --tensor-parallel-size 1SGLang
sglang serve --model-path nvidia/NVIDIA-Nemotron-3-Nano-4B-BF16 --tp 1Generic example, not from the model's own docs: adjust flags (quantization, context length, parallelism) for your setup.
NVIDIA-Nemotron-3-Nano-30B-A3B-BF16 (31.6B)
Needs about 70.6 GB of VRAM at BF16. Cheapest live fit: A100 at $1.21/hr.
vLLM
vllm serve nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16 --tensor-parallel-size 1SGLang
sglang serve --model-path nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16 --tp 1Generic example, not from the model's own docs: adjust flags (quantization, context length, parallelism) for your setup.
NVIDIA-Nemotron-3-Super-120B-A12B-BF16 (123.6B)
Needs about 276 GB of VRAM at BF16. Cheapest live fit: 6× RTX A6000 at $2.18/hr.
vLLM
vllm serve nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16 --tensor-parallel-size 6SGLang
sglang serve --model-path nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16 --tp 6Generic example, not from the model's own docs: adjust flags (quantization, context length, parallelism) for your setup.
NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 (560.5B)
Needs about 1253 GB of VRAM at BF16. No live GPU fit right now.
vLLM
vllm serve nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16SGLang
sglang serve --model-path nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16Generic example, not from the model's own docs: adjust flags (quantization, context length, parallelism) for your setup.
Launch on Aquanode
Aquanode sells GPU pods billed per second, not a hosted inference API. Rent a pod sized to the model above, open a terminal on it or save a command as a startup script, then connect to the endpoint it serves.
Sources
Updated 2026-10-07.