The NVIDIA GH200 Grace Hopper Superchip is a single module that pairs a 72-core Arm-based Grace CPU with a Hopper GPU over a 900 GB/s coherent link called NVLink-C2C. The GPU gets 96 GB of HBM3 or 144 GB of HBM3e, and the CPU adds up to 480 GB of LPDDR5X that the GPU can address directly, which is why GH200 is the Hopper-generation answer for models and datasets that do not fit in GPU memory alone.
This guide covers:
- What a Grace Hopper Superchip is and how it differs from an H100 or H200 in a normal server
- The spec table, with the HBM3 and HBM3e variants side by side
- How NVLink-C2C and unified memory change what you can run
- NVIDIA's published performance claims and MLPerf results, labelled as the vendor's
- Power, cooling and software needs, and when to pick GH200 over H100 or H200
TL;DR
- GH200 is one Grace CPU plus one Hopper GPU in a single package. The GPU is the same Hopper architecture as the H100 and H200; the difference is the CPU attached to it.
- NVIDIA's tuning guide lists 96 GB HBM3 (up to 4 TB/s) or 144 GB HBM3e (up to 4.9 TB/s) on the GPU, plus up to 480 GB LPDDR5X on the CPU side.
- NVLink-C2C runs at up to 900 GB/s total, which NVIDIA says is 7x the bandwidth of x16 PCIe Gen 5. That is what lets the GPU treat CPU memory as an extension of its own.
- Verdict: pick GH200 when your working set is bigger than 80 to 141 GB (large KV caches, retrieval indexes, graph data, simulation state) and you can run on Arm. Pick an H100 or H200 when the model fits in GPU memory and you want the broadest software compatibility.
For the wider picture of datacenter accelerators, see our datacenter GPU overview.
What is the NVIDIA GH200?
A conventional GPU server has an x86 CPU and one or more GPUs joined by PCIe. Data moves between CPU memory and GPU memory across that PCIe link, and the programmer or the framework decides what lives where.
GH200 removes that boundary. NVIDIA's architecture write-up describes a Grace CPU with up to 72 Arm Neoverse V2 cores and a Hopper GPU sitting on one module, joined by NVLink-C2C, a memory-coherent chip-to-chip link. Because the link is coherent, CPU threads and GPU threads can access both memory pools without explicit copies. NVIDIA also states that Address Translation Services let the CPU and GPU share a single per-process page table, so all system-allocated memory is reachable from either processor.
In practice that means a GPU kernel can read a tensor that lives in the CPU's LPDDR5X. It is slower than reading HBM, but it is far faster than the same access over PCIe, and it removes the code you would otherwise write to stage data in and out.
GH200 was announced in 2023 and has been shipping in HBM3 and HBM3e variants. On its product page NVIDIA describes the GH200 as currently available, and groups it with the GH200 NVL2 (two superchips linked by NVLink) and the NVIDIA MGX reference design.
GH200 specs
NVIDIA's Grace performance tuning guide gives the single-superchip numbers. The H100 SXM and H200 SXM columns come from NVIDIA's H100 and H200 product pages, for comparison.
| Spec | GH200 (HBM3) | GH200 (HBM3e) | H100 SXM | H200 SXM |
|---|---|---|---|---|
| GPU memory | 96 GB HBM3 | 144 GB HBM3e | 80 GB | 141 GB |
| GPU memory bandwidth | Up to 4 TB/s | Up to 4.9 TB/s | 3.35 TB/s | 4.8 TB/s |
| CPU | 72 Arm Neoverse V2 cores | 72 Arm Neoverse V2 cores | Separate host CPU | Separate host CPU |
| CPU memory | Up to 480 GB LPDDR5X | Up to 480 GB LPDDR5X | Host-defined | Host-defined |
| CPU memory bandwidth | Up to 500 GB/s | Up to 500 GB/s | Host-defined | Host-defined |
| CPU to GPU link | NVLink-C2C, up to 900 GB/s | NVLink-C2C, up to 900 GB/s | PCIe Gen 5, 128 GB/s | PCIe Gen 5, 128 GB/s |
Two notes on reading this table. First, NVIDIA's product page states a combined "fast memory" figure of up to 624 GB for the superchip, which matches 480 GB of LPDDR5X plus 144 GB of HBM3e. Second, NVIDIA's own pages are not consistent on small details: the Grace Hopper architecture post quotes up to 512 GB of LPDDR5X at up to 546 GB/s, while the tuning guide says up to 480 GB and up to 500 GB/s. We use the tuning guide figures because they are the more recent and match the shipping part; check the datasheet for the exact SKU you deploy.
The Hopper GPU in a GH200 has the same tensor core generation as the H100 and H200, including the FP8 Transformer Engine (see our Transformer Engine guide). NVIDIA's GH200 page does not publish a separate peak TFLOPS table for the GPU, so do not assume the GPU runs at the H100 SXM peak figures inside every power configuration.
Architecture and form factors
NVLink-C2C and unified memory
NVLink-C2C is the feature that defines the product. NVIDIA lists it at up to 900 GB/s total, 450 GB/s in each direction, and calls it 7x faster than PCIe Gen 5 x16. Three consequences follow:
- One address space. The GPU can load from CPU memory with ordinary pointers when memory is allocated through the system allocator, thanks to the shared page table.
- Oversubscription is usable. A model whose weights or KV cache exceed HBM can spill to LPDDR5X and still run, at the cost of lower bandwidth on the spilled part. LPDDR5X bandwidth (up to 500 GB/s per the tuning guide) is roughly an order of magnitude below HBM3e, so spilling is a capacity tool, not a free lunch. See our glossary on the KV cache for why long contexts are the usual trigger.
- Less copy code. NVIDIA claims up to 36x faster data processing on its product page by removing CPU to GPU memory copies. That is the vendor's number for a data processing workload, not a general speedup.
Single superchip, NVL2 and larger systems
- GH200 (single). One Grace plus one Hopper. This is the unit most people mean.
- GH200 NVL2. Two superchips connected by NVLink. NVIDIA's page lists up to 288 GB of high-bandwidth memory, 10 TB/s of memory bandwidth and 1.2 TB of fast memory for the pair, and the tuning guide gives 144 Arm cores and up to 960 GB of LPDDR5X.
- MGX. NVIDIA's modular server reference design, which lets OEMs build GH200 systems with BlueField-3 DPUs and their own I/O.
- NVLink Switch systems. NVIDIA's architecture post describes an NVLink Switch System connecting up to 256 Grace Hopper Superchips over NVLink 4, which is the basis of the DGX GH200 design. NVIDIA's DGX GH200 page describes it as a turnkey system with shared memory across the interconnected superchips but does not publish a superchip count or memory total on that page. For how DGX, HGX and rack-scale NVL systems relate, see HGX vs DGX vs NVL72.
If you want the generation that follows, the rack-scale Grace Blackwell design is covered in our GB200 NVL72 guide. It keeps the Grace-plus-GPU idea and scales it to a rack.
Performance: what NVIDIA has published
Every figure below is NVIDIA's claim or an MLPerf result, with the conditions it states. We have not benchmarked GH200 ourselves for this post.
NVIDIA product page claims (vendor's claims, no independent verification):
- Up to 10x higher performance for applications running terabytes of data.
- Up to 36x faster data processing by removing CPU to GPU memory copies.
- 30x faster embedding generation for retrieval-augmented generation.
- Up to 8x faster graph neural network training than an H100 PCIe GPU.
- GH200 NVL2 versus a single H100: up to 3.5x the GPU memory capacity and 3x the bandwidth.
These are best-case figures, chosen for workloads where memory movement dominates. They tell you where GH200 is designed to win, not what a typical LLM run will see.
MLPerf Inference v4.1 (NVIDIA's write-up of the MLCommons results, published September 24, 2024):
- Per accelerator, GH200 delivered up to 1.4x the performance of H100 on Llama 2 70B, 1.2x on Mixtral 8x7B and 1.3x on DLRMv2 99%.
- On Llama 2 70B, GH200's server-scenario performance stayed within 5% of its offline performance, while the best CPU-only x86 submission dropped 55% from offline to server.
- Against a two-socket Xeon 8592+ system, a single GH200 delivered up to 22x higher throughput on GPT-J. There were no CPU-only submissions for Llama 2 70B or Mixtral 8x7B.
MLPerf results are measured by the submitter under MLCommons rules, so the conditions are public and auditable, but the workloads are fixed benchmarks. Treat them as a reasonable proxy for LLM inference with a large memory pool, not a prediction for your model.
Infrastructure needs
Power and cooling. NVIDIA's Grace architecture post says the HGX Grace Hopper platform "can be air or liquid cooled and has up to 1,000W TDP" for the module. Plan on a module that draws considerably more than a PCIe card and budget rack power accordingly. Per-system power depends on how many superchips the OEM puts in the chassis.
CPU architecture. The host CPU is Arm (aarch64), not x86. This is the single most common surprise. Container images, Python wheels and CUDA libraries need arm64 builds. Check that every dependency you rely on has an arm64 build before you commit to the platform; one long-tail package without it will block you.
Networking. Single-node jobs need nothing special. For multi-node training, plan on the same fabric you would use for H100 clusters. Our what is NVLink guide explains when scale-up links matter and when network bandwidth does.
Software. CUDA, cuDNN and the usual frameworks work. The one change to understand is memory allocation: to use the shared page table, allocate with the system allocator rather than only cudaMalloc, and read NVIDIA's Grace performance tuning guide for memory placement guidance.
When to choose GH200
Choose GH200 when:
- Your working set exceeds GPU memory. Large retrieval indexes, embedding tables, big graphs, long-context KV caches and simulation state are the clearest wins, because NVLink-C2C makes the spill tolerable.
- CPU and GPU phases alternate. Data preparation on the CPU followed by GPU compute, repeated in a loop, benefits from not copying across PCIe.
- You want more HBM per GPU than an H100. The HBM3e variant offers 144 GB against 80 GB on the H100 SXM, the same class as the H200.
- Your model and KV cache fit in GPU memory, so the CPU memory is dead weight.
- You depend on x86-only software.
- You need 8-GPU nodes with NVSwitch inside one chassis. The HGX H100 and H200 boards are built for that; GH200 is a one-GPU-per-superchip design that scales through NVLink Switch systems or the network.
Choose Blackwell when you need much more compute or memory bandwidth per GPU: see the B200 guide, and for how the generations stack up, Rubin vs Blackwell vs Hopper.
Cost: how to think about tokens per dollar
We do not quote a rental price here, because it changes by the hour; the live box below shows the current figure. What you can do is turn a measured throughput into a cost:
- Tokens per GPU-hour equals your measured tokens per second times 3,600.
- Cost per million tokens equals the hourly price divided by tokens per GPU-hour, times 1,000,000.
Use throughput you measured on your own model, or a published MLPerf per-accelerator figure for a comparable workload, and multiply by the live hourly price shown below. GH200 tends to look best on this arithmetic when the alternative is renting extra GPUs only to hold memory, because one GH200 can replace that capacity with LPDDR5X.
Rent today
Aquanode manages and optimizes GPUs for training and inference workloads, and you can rent the GPUs in the box below on demand. Prices are live, so the table shows today's from-price per GPU-hour.
See the GH200 page for specs and availability, and /pricing for how billing works.
What's next
GH200 is the first generation of NVIDIA's Arm CPU plus GPU superchip line. The Blackwell generation continues it with the GB200 Grace Blackwell Superchip in rack-scale NVL72 systems, and NVIDIA has described Vera Rubin as the generation after that. For the status of those parts, read our Rubin guide rather than treating any ship date as settled. Compare the whole field on the GPU index.
FAQ
What is the difference between GH200 and H100?
The GPU architecture is the same (Hopper). GH200 adds a Grace Arm CPU on the same module, connected by 900 GB/s NVLink-C2C, plus up to 480 GB of LPDDR5X the GPU can address. An H100 is a GPU only and relies on a separate host CPU over PCIe.
How much memory does a GH200 have?
NVIDIA lists 96 GB of HBM3 or 144 GB of HBM3e on the GPU, and up to 480 GB of LPDDR5X on the CPU. NVIDIA's product page quotes up to 624 GB of combined fast memory per superchip.
Is GH200 faster than H100?
For memory-bound inference, NVIDIA's MLPerf Inference v4.1 submission reports up to 1.4x per accelerator on Llama 2 70B. For compute-bound work that fits in GPU memory, do not expect a large gap, since the Hopper GPU is the same generation.
Does GH200 run x86 software?
No. The CPU is Arm (aarch64), so you need arm64 builds of your containers and dependencies. The GPU side runs standard CUDA code.
What is GH200 NVL2?
Two GH200 superchips connected by NVLink. NVIDIA lists up to 288 GB of HBM3e and 10 TB/s of memory bandwidth for the pair, and 1.2 TB of fast memory.
Is GH200 good for training?
It can be, especially when the model or optimizer state exceeds GPU memory. For large multi-node training, check networking first, since GH200 is one GPU per superchip. For raw training throughput per GPU, the Blackwell generation is the better target.
Sources
- NVIDIA GH200 Grace Hopper Superchip product page: https://www.nvidia.com/en-us/data-center/grace-hopper-superchip/
- NVIDIA technical blog, Grace Hopper Superchip architecture in depth: https://developer.nvidia.com/blog/nvidia-grace-hopper-superchip-architecture-in-depth/
- NVIDIA Grace Performance Tuning Guide (spec table): https://docs.nvidia.com/dccpu/grace-perf-tuning-guide/index.html
- NVIDIA DGX GH200 page: https://www.nvidia.com/en-in/data-center/dgx-gh200/
- NVIDIA press release, GH200 with HBM3e (August 8, 2023): https://nvidianews.nvidia.com/news/gh200-grace-hopper-superchip-with-hbm3e-memory
- NVIDIA technical blog, GH200 in MLPerf Inference v4.1 (September 24, 2024): https://developer.nvidia.com/blog/nvidia-gh200-grace-hopper-superchip-delivers-outstanding-performance-in-mlperf-inference-v4-1/
- NVIDIA H100 product page (H100 SXM comparison figures): https://www.nvidia.com/en-us/data-center/h100/
- NVIDIA H200 product page (H200 SXM comparison figures): https://www.nvidia.com/en-us/data-center/h200/