NVIDIA GH200 GPU: Specs, VRAM, Price & Benchmarks (2026)

The GH200 is a real GPU (launched 2023), but no provider on Aquanode is listing it for rental right now, so there is no live hourly price to quote. Availability changes as providers add and retire hardware; the specs, VRAM and price context below still apply.

How much VRAM does the GH200 have?

The GH200 has 141GB HBM3e (GPU) + 480GB LPDDR5X (Grace CPU), 624GB unified per Superchip, with 4.8 TB/s (GPU HBM3e); 900 GB/s NVLink-C2C to the CPU of peak memory bandwidth.

What can the GH200 run?

Popular open models from small to frontier scale, with the memory each needs and how many GH200 cards (141GB HBM3e (GPU) + 480GB LPDDR5X (Grace CPU), 624GB unified per Superchip each) that takes.

ModelAs publishedFP8INT4
Qwen/Qwen3-8B 8.2BBF16: not supportedFP8: not supportedINT4: not supported
Qwen/Qwen2.5-14B-Instruct 14.8BBF16: not supportedFP8: not supportedINT4: not supported
Qwen/Qwen3-32B 32.8BBF16: not supportedFP8: not supportedINT4: not supported
Qwen/Qwen-72B 72.3BBF16: not supportedFP8: not supportedINT4: not supported
MiniMaxAI/MiniMax-M2.7 228.7BFP8: not supported–INT4: not supported
deepseek-ai/DeepSeek-R1 684.5BFP8: not supported–INT4: not supported

Estimates: weights at the stated precision plus a flat 20% for KV cache and overhead, at a moderate context length. A dash means the precision is not offered for that model (it is already published at that size). INT4 needs a published quantized checkpoint. Open any model for a per-GPU breakdown.

GH200 specs

Architecturelaunched 2023
VRAM141GB HBM3e (GPU) + 480GB LPDDR5X (Grace CPU), 624GB unified per Superchip
Memory bandwidth4.8 TB/s (GPU HBM3e); 900 GB/s NVLink-C2C to the CPU
InterconnectNVLink-C2C, 900 GB/s CPU-GPU coherent link
TDPConfigurable 450-1,000W per Superchip (Grace CPU + memory alone: up to 250W)
Form factorSuperchip module (rack-scale; also sold as a single-node MGX server building block)

Specs sourced from the vendor's public product page. See the source.

GPU Glossary: What is VRAM?, HBM, Tensor Cores, CUDA Cores, TFLOPS, NVLink vs PCIe

Good for

Pairs an H200-class Hopper GPU (same 141GB HBM3e and compute) with a coherent 480GB Grace CPU memory pool over a 900 GB/s NVLink-C2C link, so the CPU and GPU share one address space: useful for workloads with a large CPU-side data pipeline (graph analytics, large embedding tables) feeding GPU compute without a PCIe hop. Aquanode does not list per-GPU GH200 rental; the live H200 pricing above is the closest per-GPU-hour equivalent using the same Hopper GPU.

Not good for

Sold as a Superchip module or a complete server, not as an individually rentable GPU on this or most other GPU marketplaces, so it is not a drop-in swap for an H100/H200 rental despite sharing the same GPU die.

Related guides

Other models in the same generation. The full list is in the GPU index.

Get notified when the price drops

GPU supply moves hourly. Tell us what you're waiting for and we'll email you when a matching offer appears across any provider we track.

One email per matching alert. Unsubscribe any time.

Submit the job. Everything after that is ours.

Sign up in 60 seconds. Pay for the GPU minutes you actually use.

© 2026 Aquanode. All rights reserved.

All trademarks, logos and brand names are the property of their respective owners.