NVIDIA H100 NVL GPU: Specs, VRAM, Price & Benchmarks (2026)

The H100 NVL is a real GPU (launched 2023), but no provider on Aquanode is listing it for rental right now, so there is no live hourly price to quote. Availability changes as providers add and retire hardware; the specs, VRAM and price context below still apply.

How much VRAM does the H100 NVL have?

The H100 NVL has 94GB HBM3, with 3.9 TB/s of peak memory bandwidth.

H100 NVL VRAM calculator: check which models fit in its memory at each precision.

All models that fit in 80 GB: the open models whose weights and overhead fit, at native, FP8 and INT4 precision.

What can the H100 NVL run?

Popular open models from small to frontier scale, with the memory each needs and how many H100 NVL cards (94GB HBM3 each) that takes.

ModelAs publishedFP8INT4
Qwen/Qwen3-8B 8.2BBF16: ~18.3 GB, 1 GPUFP8: ~9.2 GB, 1 GPUINT4: ~4.6 GB, 1 GPU
Qwen/Qwen2.5-14B-Instruct 14.8BBF16: ~33 GB, 1 GPUFP8: ~16.5 GB, 1 GPUINT4: ~8.3 GB, 1 GPU
Qwen/Qwen3-32B 32.8BBF16: ~73.2 GB, 1 GPUFP8: ~36.6 GB, 1 GPUINT4: ~18.3 GB, 1 GPU
Qwen/Qwen-72B 72.3BBF16: ~162 GB, 2 GPUsFP8: ~80.8 GB, 1 GPUINT4: ~40.4 GB, 1 GPU
MiniMaxAI/MiniMax-M2.7 228.7BFP8: ~256 GB, 3 GPUs–INT4: ~128 GB, 2 GPUs
deepseek-ai/DeepSeek-R1 684.5BFP8: ~765 GB, 9 GPUs–INT4: ~383 GB, 5 GPUs

Estimates: weights at the stated precision plus a flat 20% for KV cache and overhead, at a moderate context length. A dash means the precision is not offered for that model (it is already published at that size). INT4 needs a published quantized checkpoint. Open any model for a per-GPU breakdown, or use the H100 NVL VRAM calculator.

H100 NVL specs

Architecturelaunched 2023
VRAM94GB HBM3
Memory bandwidth3.9 TB/s
FP16 / BF16 tensor throughput835.5 TFLOPS (peak, dense)
FP8 tensor throughput1,670.5 TFLOPS (peak, dense)
InterconnectNVLink, 600 GB/s
TDP350-400W (configurable)
Form factorPCIe, dual-slot, air-cooled

Specs sourced from the vendor's public product page. See the source.

GPU Glossary: What is VRAM?, HBM, Tensor Cores, CUDA Cores, TFLOPS, NVLink vs PCIe

Good for

An H100 in a PCIe dual-slot card with 94GB instead of 80GB and higher memory bandwidth than the H100 PCIe, NVIDIA pairs two of them for 188GB of HBM3 in total. It is aimed at LLM inference: a 70B model fits in INT4 on one card, and a pair holds it in FP16.

Not good for

Its NVLink runs at 600GB/s against 900GB/s on the SXM H100, and at 350-400W it has a much lower power envelope than the 700W SXM part, which caps sustained throughput. 94GB is still not enough for a 70B model in FP16 on one card.

Related guides

Other models in the same generation, or see how it compares in the GPU benchmarks and specs table. The full list is in the GPU index.

Related reading: H100 and H200: SXM vs NVL vs PCIe.

Get notified when the price drops

GPU supply moves hourly. Tell us what you're waiting for and we'll email you when a matching offer appears across any provider we track.

One email per matching alert. Unsubscribe any time.

Submit the job. Everything after that is ours.

Sign up in 60 seconds. Pay for the GPU minutes you actually use.

© 2026 Aquanode. All rights reserved.

All trademarks, logos and brand names are the property of their respective owners.