NVIDIA GH200 GPU: Specs, VRAM, Price & Benchmarks (2026)
The GH200 is a real GPU (launched 2023), but no provider on Aquanode is listing it for rental right now, so there is no live hourly price to quote. Availability changes as providers add and retire hardware; the specs, VRAM and price context below still apply.
How much VRAM does the GH200 have?
The GH200 has 141GB HBM3e (GPU) + 480GB LPDDR5X (Grace CPU), 624GB unified per Superchip, with 4.8 TB/s (GPU HBM3e); 900 GB/s NVLink-C2C to the CPU of peak memory bandwidth.
What can the GH200 run?
Popular open models from small to frontier scale, with the memory each needs and how many GH200 cards (141GB HBM3e (GPU) + 480GB LPDDR5X (Grace CPU), 624GB unified per Superchip each) that takes.
| Model | As published | FP8 | INT4 |
|---|---|---|---|
| Qwen/Qwen3-8B 8.2B | BF16: not supported | FP8: not supported | INT4: not supported |
| Qwen/Qwen2.5-14B-Instruct 14.8B | BF16: not supported | FP8: not supported | INT4: not supported |
| Qwen/Qwen3-32B 32.8B | BF16: not supported | FP8: not supported | INT4: not supported |
| Qwen/Qwen-72B 72.3B | BF16: not supported | FP8: not supported | INT4: not supported |
| MiniMaxAI/MiniMax-M2.7 228.7B | FP8: not supported | – | INT4: not supported |
| deepseek-ai/DeepSeek-R1 684.5B | FP8: not supported | – | INT4: not supported |
Estimates: weights at the stated precision plus a flat 20% for KV cache and overhead, at a moderate context length. A dash means the precision is not offered for that model (it is already published at that size). INT4 needs a published quantized checkpoint. Open any model for a per-GPU breakdown.
GH200 specs
| Architecture | launched 2023 |
| VRAM | 141GB HBM3e (GPU) + 480GB LPDDR5X (Grace CPU), 624GB unified per Superchip |
| Memory bandwidth | 4.8 TB/s (GPU HBM3e); 900 GB/s NVLink-C2C to the CPU |
| Interconnect | NVLink-C2C, 900 GB/s CPU-GPU coherent link |
| TDP | Configurable 450-1,000W per Superchip (Grace CPU + memory alone: up to 250W) |
| Form factor | Superchip module (rack-scale; also sold as a single-node MGX server building block) |
Specs sourced from the vendor's public product page. See the source.
GPU Glossary: What is VRAM?, HBM, Tensor Cores, CUDA Cores, TFLOPS, NVLink vs PCIe
Good for
Pairs an H200-class Hopper GPU (same 141GB HBM3e and compute) with a coherent 480GB Grace CPU memory pool over a 900 GB/s NVLink-C2C link, so the CPU and GPU share one address space: useful for workloads with a large CPU-side data pipeline (graph analytics, large embedding tables) feeding GPU compute without a PCIe hop. Aquanode does not list per-GPU GH200 rental; the live H200 pricing above is the closest per-GPU-hour equivalent using the same Hopper GPU.
Not good for
Sold as a Superchip module or a complete server, not as an individually rentable GPU on this or most other GPU marketplaces, so it is not a drop-in swap for an H100/H200 rental despite sharing the same GPU die.
Related guides
Other models in the same generation. The full list is in the GPU index.
Get notified when the price drops
GPU supply moves hourly. Tell us what you're waiting for and we'll email you when a matching offer appears across any provider we track.