Nemotron 3 models
16 Nemotron 3 models on Hugging Face, from 638M to 560.5B parameters. At the precision each one is published in, the smallest needs about 1.7 GB of VRAM (NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4-DSpark, cheapest live fit: RTX 5060 Ti) and the largest about 1253 GB (NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16). The cheapest way to run NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4-DSpark is $0.110/hr.
Nemotron 3 models
VRAM is for the precision the model is published in, with the same overhead and an 8,192-token context assumed on every page; see the methodology. The fit shown is the lowest-priced single GPU type that holds the model at that precision, or the lowest-priced multi-GPU set (up to 8) when none does.
More from NVIDIA
- Parakeet (7 models)
- Cosmos (6 models)
- Nemotron Nano (4 models)
- Nemotron (5 models)
- Nemotron Labs (4 models)
- Nemotron-H (3 models)