How to fine-tune Nemotron 3 Nano (LoRA, QLoRA)

Base versus post-trained checkpoints, chat template, NeMo guidance and license terms for fine-tuning NVIDIA Nemotron 3 Nano.

This guide covers adapting the Nemotron 3 family with LoRA or QLoRA. Estimate cost with the training cost calculator and see fine-tuning for the overview.

Which checkpoint

  • The Base checkpoint (NVIDIA-Nemotron-3-Nano-30B-A3B-Base-BF16) is trained with next-token prediction. Its card calls it "a good starting point for instruction fine-tuning" for developers building instruction-following models.
  • The main BF16 checkpoint is the post-trained one with reasoning and tool calling. Start from it to keep that behaviour.

Training framework

The post-trained card recommends the NeMo Framework for fine-tuning. The cards give no LoRA target-module list, so inspect the module names yourself. The model is a Mamba2-Transformer hybrid MoE (23 MoE, 23 Mamba-2 and 6 attention layers), so confirm your trainer supports those layers and the custom code (trust_remote_code).

Chat template

Use the model's template. Reasoning is controlled by enable_thinking in apply_chat_template() and is on by default. If you train a non-reasoning variant, build your data with enable_thinking=False so the format matches inference.

License

Both checkpoints use the NVIDIA Nemotron Open Model License. Per the license text:

  • You may create and distribute derivative works, and commercial use is permitted.
  • Redistribution requires a copy of the license, retained notices, a NOTICE file, and the statement "Licensed by NVIDIA Corporation under the NVIDIA Nemotron Model License."
  • Patent or copyright litigation against NVIDIA over the work terminates your license.
  • NVIDIA does not claim ownership of outputs.

Memory to fine-tune Nemotron 3, by size

Model state in GB before activations, for a full fine-tune, LoRA and QLoRA, with the cheapest live GPU set that has that much memory. One model per size.

ModelParametersFull fine-tuneLoRAQLoRACompute per 1B tokens
NVIDIA-Nemotron-3-Nano-4B-BF164.0B59.2 GBA100 · $1.21/hr7.4 GBRTX 4070 Super · $0.121/hr1.9 GBRTX 4070 Super · $0.121/hr2.4 × 10^19 FLOPs
NVIDIA-Nemotron-3-Nano-30B-A3B-BF1631.6B471 GBRTX PRO 6000 × 5 · $6.88/hr58.8 GBA100 · $1.21/hr15.2 GBRTX A5000 · $0.176/hr1.9 × 10^20 FLOPs
NVIDIA-Nemotron-3-Super-120B-A12B-BF16123.6B1842 GBNo live fit230 GBRTX A6000 × 5 · $1.81/hr59.4 GBA100 · $1.21/hr7.4 × 10^20 FLOPs
NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16560.5B8352 GBNo live fit1044 GBNo live fit269 GBRTX A6000 × 6 · $2.18/hr3.4 × 10^21 FLOPs

Full fine-tune counts 16 bytes per parameter (mixed-precision Adam, as counted in the ZeRO paper). LoRA keeps the frozen base in BF16 at 2 bytes per parameter. QLoRA stores the base in 4-bit NormalFloat with double quantization at 4.127 bits per parameter (QLoRA paper). Adapters and activations are not counted: activations depend on your batch size and sequence length, so leave headroom. Compute is the training-cost calculator's 6 × parameters × tokens for a dense model; mixture-of-experts models use fewer. These are estimates from formulas, not measurements of a run.

Estimate a full run

Turn the compute column into time and cost with the training cost calculator. For the methods themselves, read LoRA and QLoRA, and see how fine-tuning works on Aquanode at fine-tuning.

Sources

Updated 2026-10-07.

Related

Submit the job. Everything after that is ours.

Sign up in 60 seconds. Pay for the GPU minutes you actually use.

© 2026 Aquanode. All rights reserved.

All trademarks, logos and brand names are the property of their respective owners.