The NVIDIA GB200 NVL72 is a liquid-cooled rack that links 72 Blackwell GPUs and 36 Grace CPUs into a single NVLink domain, so the whole rack behaves like one very large GPU. NVIDIA lists 13.4 TB of HBM3E, 130 TB/s of NVLink bandwidth and 1,440 PFLOPS of sparse FP4 for the rack, and "GB200" on its own usually means the Grace Blackwell Superchip inside it: one Grace CPU plus two Blackwell GPUs.
This guide covers:
- What GB200 and GB200 NVL72 actually are, and how the names differ
- The rack and Superchip specs, with per-GPU figures worked out from them
- NVIDIA's published performance claims and what each one is measured against
- Power, cooling and networking you need before a rack can run
- Who should use a rack-scale system and who is better off with an 8-GPU B200 server
TL;DR
- GB200 is the Grace Blackwell Superchip: 1 Grace CPU and 2 Blackwell GPUs joined by NVLink-C2C. GB200 NVL72 is 36 of those in one rack, giving 72 GPUs in one NVLink domain.
- The point of the rack is the interconnect. Every GPU talks to every other GPU at 1.8 TB/s through NVLink switches, which matters for trillion-parameter models and large mixture-of-experts inference.
- Supermicro's rack listing puts a rack at 132 kW and requires direct liquid cooling. This is a facility project, not a server you drop into an air-cooled room.
- For models that fit in 8 GPUs, an HGX B200 server is simpler and cheaper to run. Rack-scale pays off when one model has to span many GPUs with fast links.
- Verdict: GB200 NVL72 is the right tool for frontier-scale serving and training. For most teams, a B200 or H200 node is the practical choice today.
GB200 vs GB200 NVL72: what the names mean
NVIDIA uses several related names, and search results mix them up.
- Blackwell GPU (B200): the GPU itself. NVIDIA describes it as two reticle-limited dies joined by a 10 TB/s chip-to-chip interconnect, 208 billion transistors, on a custom TSMC 4NP process. See our B200 guide.
- GB200 Grace Blackwell Superchip: 1 Grace CPU (72 Arm Neoverse V2 cores) and 2 Blackwell GPUs on one board, linked by NVLink-C2C at 900 GB/s of bidirectional bandwidth.
- GB200 NVL72: the rack. 18 compute trays, each holding two Grace CPUs and four Blackwell GPUs, plus nine NVLink switch trays. "72" is the GPU count.
- DGX GB200: NVIDIA's own branded system built on the same 36-Superchip, 72-GPU rack, sold with Mission Control software and support. NVIDIA's DGX GB200 page links to the NVL72 product page but does not itself use the NVL72 name.
So when someone asks for "a GB200," ask whether they mean a Superchip, a compute tray, or the full rack. We cover how DGX, HGX and NVL72 relate in HGX vs DGX vs NVL72.
GB200 NVL72 specs
NVIDIA's product page gives the rack and Superchip figures below. FP4 and FP8 numbers are NVIDIA's with-sparsity figures unless marked dense.
| Spec | GB200 NVL72 (rack) | GB200 Superchip (1 CPU, 2 GPUs) |
|---|---|---|
| GPUs / CPUs | 72 Blackwell / 36 Grace | 2 Blackwell / 1 Grace |
| CPU cores | 2,592 Arm Neoverse V2 | 72 Arm Neoverse V2 |
| GPU memory | 13.4 TB HBM3E, 576 TB/s | 372 GB HBM3E, 16 TB/s |
| CPU memory | 17 TB LPDDR5X, 14 TB/s | up to 480 GB LPDDR5X, up to 512 GB/s |
| NVLink bandwidth | 130 TB/s | 3.6 TB/s |
| FP4 Tensor Core | 1,440 PFLOPS sparse, 720 dense | 40 PFLOPS sparse, 20 dense |
| FP8/FP6 Tensor Core | 720 PFLOPS sparse | 20 PFLOPS sparse |
| FP16/BF16 Tensor Core | 360 PFLOPS | 10 PFLOPS |
| FP64 | 2,880 TFLOPS | 80 TFLOPS |
Dividing the rack figures by 72 gives per-GPU numbers (our arithmetic, not an NVIDIA table): about 186 GB of HBM3E per GPU, about 20 PFLOPS of sparse FP4 per GPU, and 8 TB/s of memory bandwidth per GPU. NVIDIA's technical blog also cites 1.8 TB/s of bidirectional NVLink bandwidth per GPU, about 14 times PCIe Gen5.
The 17 TB of LPDDR5X attached to the Grace CPUs is part of the same coherent address space, which is why NVIDIA also quotes 30 TB of "fast memory" for the rack. Treat that CPU memory as a slower tier: 14 TB/s across the rack is far below the 576 TB/s of HBM3E. For the glossary background, see HBM and NVSwitch.
Architecture: why a rack, not a server
An eight-GPU server like the HGX B200 gives you a 1.8 TB/s NVLink fabric across 8 GPUs. Beyond 8 GPUs you cross to InfiniBand or Ethernet, which is much slower per link. The NVL72 extends the NVLink fabric across a rack.
NVIDIA's technical blog describes the build: 18 compute trays, nine NVLink switch trays with 144 NVLink ports each, wired so that all 18 NVLink ports on each of the 72 GPUs reach the switches. The result is a 72-GPU, all-to-all NVLink domain.
Why that matters in practice:
- Tensor and expert parallelism. Splitting one model across many GPUs sends activations between them every layer. Mixture-of-experts models route tokens to experts spread across GPUs, which is communication heavy. Keeping that traffic on NVLink instead of the network is the main reason NVIDIA built this system. See tensor parallelism and mixture of experts in the glossary.
- Large KV caches. With 13.4 TB of HBM3E in one domain, long-context serving can pool memory across GPUs.
- Grace CPU coherence. Each GPU pair shares a coherent link to a Grace CPU, which helps offloading and data-heavy pipelines.
For the NVLink technology itself, read what is NVLink.
Performance: what NVIDIA claims
Every figure here is NVIDIA's own, and NVIDIA labels most of them as projections that are subject to change. We have not measured any of them. The baselines matter, so each is spelled out.
- Real-time inference, 30x. NVIDIA's product page claims "30x faster real-time trillion-parameter large language model (LLM) inference" against HGX H100 scaled over InfiniBand. Conditions: 50 ms token-to-token latency, 5 s first-token latency, 32,768 input and 1,024 output tokens, on a 32,768-GPU cluster. The technical blog describes the 1.8T-parameter GPT-MoE test and shows 150 tokens per second per GPU for GB200 against 3.4 for H100.
- Training, 4x. Four times faster than H100 for a 1.8T-parameter mixture-of-experts model, comparing 4,096 HGX H100 GPUs with 456 GB200 NVL72 racks, both over InfiniBand.
- Energy, 25x. NVIDIA says 25 times more performance at the same power versus H100 air-cooled infrastructure.
- MLPerf Inference v5.0 (a published, audited result). In NVIDIA's write-up of the round, GB200 NVL72 delivered up to 3.4x higher per-GPU performance on Llama 3.1 405B than an eight-GPU H200 system (2.8x offline, 3.4x server). The same post notes that per-GPU performance is not a primary MLPerf metric, and says the system-level gain was up to 30x because the rack puts 9x more GPUs in one NVLink domain.
Read these as "best case, vendor-chosen workload." The 30x number in particular compares a rack-scale NVLink system against a cluster that has to cross a network, on a model sized to show that gap.
Infrastructure: power, cooling, networking
NVIDIA's product page says the NVL72 is liquid-cooled but does not state a rack power figure. OEM listings do:
- Power. Supermicro lists "total power 132kW" for its GB200 NVL72 rack. (The same line lists eight 33 kW power units, which would add up to more than 132 kW, so confirm the redundancy breakdown with your vendor.) NVIDIA does not state a rack power figure. For comparison, NVIDIA lists about 14.3 kW maximum for a single DGX B200 server.
- Cooling. Direct liquid cooling to the GPUs and CPUs through a coolant distribution unit. Supermicro lists a 250 kW in-rack CDU as standard, a 1.3 MW in-row option, and 180 kW or 240 kW liquid-to-air options for sites without a cooling tower and water supply.
- Form factor. Supermicro's rack is a 48U, 19-inch enclosure.
- Networking. For scale-out beyond one rack, NVIDIA's DGX GB200 page lists 72 OSFP ConnectX-7 ports at 400 Gb/s InfiniBand and 36 dual-port BlueField-3 DPUs. OEM listings cite up to 400 Gb/s Quantum-2 InfiniBand or Spectrum-X Ethernet. See NVLink vs InfiniBand.
- Software. NVIDIA lists Mission Control, NVIDIA AI Enterprise and DGX OS for the DGX version.
The practical consequence: you cannot put a GB200 NVL72 in a standard air-cooled colo cage. Few facilities offer racks in the 130 kW class with liquid loops, which is why most teams reach this hardware through a provider rather than buying it.
When to choose GB200 NVL72
Choose rack-scale when:
- One model needs far more than 8 GPUs and its parallelism is communication heavy, such as very large mixture-of-experts serving with expert parallelism.
- You serve long contexts where pooled HBM across many GPUs removes a bottleneck.
- You train frontier-scale models and the scale-up fabric, not raw FLOPS, limits your step time.
Choose a smaller system when:
- Your model fits in one 8-GPU node. A B200 server or an H200 gives you the same Blackwell or Hopper GPU without the facility burden.
- You want the newest memory per GPU: see GB300 NVL72 vs GB200 NVL72 for the Blackwell Ultra rack.
- Your workload is many small independent requests. Extra NVLink does little for that.
Our side-by-side of the single-GPU generations is in H200 vs B200 vs GB200. The full chip map is in the datacenter GPU guide, and the GPU page for this chip is /gpu/nvidia-gb200.
Cost: how to think about it
We do not quote an hourly price here, because rack-scale rental prices move and the live figure sits in the box below. The way to compare is cost per token. NVIDIA's DGX B200 page cites SemiAnalysis InferenceX (Q1 2026) at about $0.02 per million tokens on GPT-OSS-120B with TensorRT-LLM for that system, which NVIDIA calls roughly 4.5 times cheaper than Hopper at $0.09 per million tokens. That is an eight-GPU Blackwell figure, not a GB200 NVL72 figure, and it is NVIDIA's citation of a third-party benchmark.
To estimate your own cost: take your measured tokens per second per GPU, multiply by 3,600 for tokens per GPU-hour, then multiply the live hourly price below by your GPU count and divide by the total tokens. Measure on your own model, because the best case for a rack depends heavily on model size and latency target.
Rent today
Aquanode manages and optimizes GPUs for training and inference workloads, and you can rent the GPUs listed below on demand. The box shows what is available right now; if a chip says "None right now," nothing is currently on offer.
What's next
NVIDIA's Blackwell Ultra rack, GB300 NVL72, keeps the 72-GPU layout and raises GPU memory to 20 TB, which we compare in GB300 NVL72 vs GB200 NVL72. After that, NVIDIA's Vera Rubin NVL72 page says the rack is "ramping into full production" and claims one-tenth the cost per million tokens against GB200 NVL72 on Kimi-K2-Thinking (vendor claim, projected). See Vera Rubin NVL72 and Rubin vs Blackwell vs Hopper.
FAQ
What is the difference between GB200 and B200?
B200 is a Blackwell GPU, usually sold in eight-GPU servers. GB200 is a Superchip that pairs one Grace CPU with two Blackwell GPUs, and GB200 NVL72 racks 36 of them together with NVLink switches.
How many GPUs are in a GB200 NVL72?
72 Blackwell GPUs and 36 Grace CPUs, in 18 compute trays with nine NVLink switch trays, according to NVIDIA.
How much power does a GB200 NVL72 draw?
NVIDIA's product page does not state it. Supermicro's listing gives 132 kW per rack, and the rack needs direct liquid cooling.
How much does a GB200 NVL72 cost to buy?
NVIDIA does not publish a list price. Systems are sold through NVIDIA and its OEM partners, so quotes vary. We do not estimate one here.
Is GB200 NVL72 the same as DGX GB200?
DGX GB200 is NVIDIA's branded system built on the same 36-Superchip, 72-GPU rack layout, sold with DGX software and support. NVIDIA's DGX GB200 page does not use the NVL72 name itself.
Is GB200 faster than H100?
NVIDIA claims up to 30x for real-time trillion-parameter inference and 4x for training at cluster scale, against HGX H100 (vendor claims, projected). In MLPerf v5.0, NVIDIA reported up to 3.4x per GPU against H200 on Llama 3.1 405B.
Sources
- NVIDIA GB200 NVL72 product page: https://www.nvidia.com/en-us/data-center/gb200-nvl72/
- NVIDIA DGX GB200 product page: https://www.nvidia.com/en-us/data-center/dgx-gb200/
- NVIDIA technical blog, GB200 NVL72 training and inference: https://developer.nvidia.com/blog/nvidia-gb200-nvl72-delivers-trillion-parameter-llm-training-and-real-time-inference/
- NVIDIA Blackwell architecture page: https://www.nvidia.com/en-us/data-center/technologies/blackwell-architecture/
- NVIDIA technical blog, MLPerf Inference v5.0: https://developer.nvidia.com/blog/nvidia-blackwell-delivers-massive-performance-leaps-in-mlperf-inference-v5-0
- NVIDIA DGX B200 product page: https://www.nvidia.com/en-us/data-center/dgx-b200/
- NVIDIA Vera Rubin NVL72 page: https://www.nvidia.com/en-us/data-center/vera-rubin-nvl72/
- NVIDIA GB300 NVL72 page: https://www.nvidia.com/en-us/data-center/gb300-nvl72/
- Supermicro SRS-GB200-NVL72 rack listing: https://www.supermicro.com/en/products/rack/srs-gb200-nvl72