Hunyuan-A13B-Instruct vs NVIDIA-Nemotron-3-Super-120B-A12B-Base-BF16
Hunyuan-A13B-Instruct (80.4B parameters) and NVIDIA-Nemotron-3-Super-120B-A12B-Base-BF16 (123.6B parameters) side by side: the memory each needs at every precision, what it costs to run on a live GPU, and the context window, KV cache and license where they are published. Numbers are computed from the models' published specs; this page does not rank quality.
Side by side
| Fact | Hunyuan-A13B-Instruct | NVIDIA-Nemotron-3-Super-120B-A12B-Base-BF16 |
|---|---|---|
| Parameters | 80.4B | 123.6B |
| Architecture | Grouped-query attention; mixture of 64 experts | Hybrid (some layers use full attention); mixture of 512 experts, 22 active per token |
| Context length | 32,768 tokens | 1,048,576 tokens |
| License | – | – |
| Published precision | BF16 | BF16 |
| VRAM needed, As published | 180 GB | 276 GB |
| VRAM needed, FP8 | 89.8 GB | 138 GB |
| VRAM needed, INT4 | 44.9 GB | 69.1 GB |
| Cheapest live fit, As published | RTX A5000 × 8 · $1.41/hr | RTX A6000 × 6 · $2.18/hr |
| Cheapest live fit, FP8 | RTX PRO 6000 · $1.38/hr | RTX 4000 SFF Ada × 7 · $1.39/hr |
| Cheapest live fit, INT4 | RTX A6000 · $0.363/hr | A100 · $1.21/hr |
| KV cache per token (16-bit) | 128 KB | 8 KB |
| KV cache at 32k tokens | 4.00 GB | 0.25 GB |
| KV cache at 128k tokens | 16.0 GB | 1.00 GB |
VRAM is the weight size at each precision times a flat 1.2 overhead; see the methodology. The FP8 and INT4 rows need a quantized checkpoint or an engine that quantizes on load. The fit is the lowest-priced single GPU type that holds the model at that precision, or the lowest-priced multi-GPU set (up to 8) when none does. KV cache is for one sequence at 16-bit, computed from each model's config where the attention layout is known.
Which to pick
- Hunyuan-A13B-Instruct needs less VRAM at its published precision (180 GB against 276 GB), so it fits on a smaller GPU.
- NVIDIA-Nemotron-3-Super-120B-A12B-Base-BF16 lists the longer context window (1,048,576 tokens against 32,768).
- NVIDIA-Nemotron-3-Super-120B-A12B-Base-BF16 caches less per sequence at 32k tokens (0.3 GB against 4.0 GB), leaving more memory for batching.
- Hunyuan-A13B-Instruct has the cheaper live GPU fit at its published precision ($1.41/hr against $2.18/hr).
These follow only from the facts in the table above. Whether either model does your task well is a separate question this page does not answer.
Keep reading
- Hunyuan-A13B-Instruct: full VRAM table and live GPU fit
- NVIDIA-Nemotron-3-Super-120B-A12B-Base-BF16: full VRAM table and live GPU fit
- The Hunyuan model series
- The Nemotron model series
- All models that fit in 192 GB
- All models that fit in 288 GB
Other comparisons
- Hunyuan-A13B-Instruct vs NVIDIA-Nemotron-3-Super-120B-A12B-BF16
- Hunyuan-A13B-Instruct vs Qwen3-Next-80B-A3B-Instruct
- Hunyuan-A13B-Instruct vs GLM-4.5-Air
- Hunyuan-A13B-Instruct vs Qwen3-Next-80B-A3B-Thinking
- NVIDIA-Nemotron-3-Super-120B-A12B-Base-BF16 vs Qwen3-Next-80B-A3B-Instruct
- NVIDIA-Nemotron-3-Super-120B-A12B-Base-BF16 vs GLM-4.5-Air
- NVIDIA-Nemotron-3-Super-120B-A12B-Base-BF16 vs Qwen3-Next-80B-A3B-Thinking
- NVIDIA-Nemotron-3-Super-120B-A12B-Base-BF16 vs GLM-4.5-Air-Base