NVIDIA H100 NVL GPU: Specs, VRAM, Price & Benchmarks (2026)
The H100 NVL is a real GPU (launched 2023), but no provider on Aquanode is listing it for rental right now, so there is no live hourly price to quote. Availability changes as providers add and retire hardware; the specs, VRAM and price context below still apply.
How much VRAM does the H100 NVL have?
The H100 NVL has 94GB HBM3, with 3.9 TB/s of peak memory bandwidth.
H100 NVL VRAM calculator: check which models fit in its memory at each precision.
All models that fit in 80 GB: the open models whose weights and overhead fit, at native, FP8 and INT4 precision.
What can the H100 NVL run?
Popular open models from small to frontier scale, with the memory each needs and how many H100 NVL cards (94GB HBM3 each) that takes.
| Model | As published | FP8 | INT4 |
|---|---|---|---|
| Qwen/Qwen3-8B 8.2B | BF16: ~18.3 GB, 1 GPU | FP8: ~9.2 GB, 1 GPU | INT4: ~4.6 GB, 1 GPU |
| Qwen/Qwen2.5-14B-Instruct 14.8B | BF16: ~33 GB, 1 GPU | FP8: ~16.5 GB, 1 GPU | INT4: ~8.3 GB, 1 GPU |
| Qwen/Qwen3-32B 32.8B | BF16: ~73.2 GB, 1 GPU | FP8: ~36.6 GB, 1 GPU | INT4: ~18.3 GB, 1 GPU |
| Qwen/Qwen-72B 72.3B | BF16: ~162 GB, 2 GPUs | FP8: ~80.8 GB, 1 GPU | INT4: ~40.4 GB, 1 GPU |
| MiniMaxAI/MiniMax-M2.7 228.7B | FP8: ~256 GB, 3 GPUs | – | INT4: ~128 GB, 2 GPUs |
| deepseek-ai/DeepSeek-R1 684.5B | FP8: ~765 GB, 9 GPUs | – | INT4: ~383 GB, 5 GPUs |
Estimates: weights at the stated precision plus a flat 20% for KV cache and overhead, at a moderate context length. A dash means the precision is not offered for that model (it is already published at that size). INT4 needs a published quantized checkpoint. Open any model for a per-GPU breakdown, or use the H100 NVL VRAM calculator.
H100 NVL specs
| Architecture | launched 2023 |
| VRAM | 94GB HBM3 |
| Memory bandwidth | 3.9 TB/s |
| FP16 / BF16 tensor throughput | 835.5 TFLOPS (peak, dense) |
| FP8 tensor throughput | 1,670.5 TFLOPS (peak, dense) |
| Interconnect | NVLink, 600 GB/s |
| TDP | 350-400W (configurable) |
| Form factor | PCIe, dual-slot, air-cooled |
Specs sourced from the vendor's public product page. See the source.
GPU Glossary: What is VRAM?, HBM, Tensor Cores, CUDA Cores, TFLOPS, NVLink vs PCIe
Good for
An H100 in a PCIe dual-slot card with 94GB instead of 80GB and higher memory bandwidth than the H100 PCIe, NVIDIA pairs two of them for 188GB of HBM3 in total. It is aimed at LLM inference: a 70B model fits in INT4 on one card, and a pair holds it in FP16.
Not good for
Its NVLink runs at 600GB/s against 900GB/s on the SXM H100, and at 350-400W it has a much lower power envelope than the 700W SXM part, which caps sustained throughput. 94GB is still not enough for a 70B model in FP16 on one card.
Related guides
Other models in the same generation, or see how it compares in the GPU benchmarks and specs table. The full list is in the GPU index.
Related reading: H100 and H200: SXM vs NVL vs PCIe.
Get notified when the price drops
GPU supply moves hourly. Tell us what you're waiting for and we'll email you when a matching offer appears across any provider we track.