GB300 NVL72 is the Blackwell Ultra upgrade of GB200 NVL72: the same 72-GPU, 36-CPU liquid-cooled rack layout, but with about 1.5 times the HBM3E (20 TB against 13.4 TB), 1.5 times the dense FP4 compute, and 800 Gb/s of network bandwidth per GPU. NVLink bandwidth stays at 130 TB/s, so the rack-scale fabric did not change; the GPUs and the memory around them did.
TL;DR
- Same rack, better GPUs. Both racks hold 72 GPUs and 36 Grace CPUs and share 130 TB/s of NVLink bandwidth. GB300 swaps in Blackwell Ultra GPUs.
- Memory is the headline. NVIDIA lists 20 TB of GPU memory for GB300 NVL72 against 13.4 TB for GB200 NVL72, which NVIDIA describes as 1.5x larger HBM3E.
- Compute gain is FP4 and attention. NVIDIA claims 1.5x more dense FP4 Tensor Core FLOPS and 2x attention-layer acceleration. FP8 and FP16 rack figures are unchanged on NVIDIA's pages.
- Some precision was traded away. NVIDIA's GB300 page lists much lower INT8 and FP64 numbers than the GB200 page does.
- Verdict: pick GB300 for long-context and large-model inference where memory limits you. GB200 is fine if you already have it or the memory fits.
Spec comparison
All figures below are from NVIDIA's product pages for each rack. FP4 is the sparse figure unless marked.
| Spec | GB200 NVL72 | GB300 NVL72 |
|---|---|---|
| GPUs | 72 Blackwell | 72 Blackwell Ultra |
| CPUs | 36 Grace (2,592 Arm Neoverse V2 cores) | 36 Grace (2,592 Arm Neoverse V2 cores) |
| GPU memory | 13.4 TB HBM3E | 20 TB |
| GPU memory bandwidth | 576 TB/s | up to 576 TB/s |
| CPU memory | 17 TB LPDDR5X, 14 TB/s | 17 TB LPDDR5X, 14 TB/s |
| Total fast memory | 30 TB | 37 TB |
| NVLink bandwidth | 130 TB/s | 130 TB/s |
| FP4 Tensor Core | 1,440 PFLOPS sparse, 720 dense | 1,440 PFLOPS sparse, 1,080 dense |
| FP8/FP6 Tensor Core | 720 PFLOPS | 720 PFLOPS |
| FP16/BF16 Tensor Core | 360 PFLOPS | 360 PFLOPS |
| INT8 Tensor Core | 720 POPS | 24 POPS |
| FP64 | 2,880 TFLOPS | 100 TFLOPS |
| Network per GPU | 400 Gb/s ConnectX-7 (DGX GB200 page) | 800 Gb/s ConnectX-8 |
| Rack power | not stated by NVIDIA | not stated by NVIDIA |
Three notes on reading that table:
- NVIDIA's GB300 NVL72 page lists FP4 as 1,440 and 1,080 PFLOPS, with a footnote marking the second figure "without sparsity"; the page does not label the two columns more directly. The 1,080 dense figure is also exactly 1.5 times the 720 dense figure on the GB200 page. The 30 TB and 37 TB totals for fast memory (GPU plus CPU memory) come from NVIDIA's DGX GB200 page and the GB300 page respectively.
- Dividing by 72 GPUs (our arithmetic) gives about 186 GB of HBM3E per GPU on GB200 and about 278 GB on GB300.
- The INT8 and FP64 drop is real on NVIDIA's pages. Blackwell Ultra is tuned for low-precision AI inference. If your workload leans on FP64 (some HPC and simulation codes), check the numbers before assuming GB300 is a straight upgrade.
What changed and why it matters
Memory per GPU
The 1.5x increase in HBM3E is the change most buyers will feel. A bigger HBM stack per GPU means larger models, longer contexts and bigger KV caches fit before you have to shard further. See HBM and KV cache in the glossary. The same step is what separates the single-GPU B300 from the B200; our B300 vs B200 post covers it per GPU.
FP4 and attention
NVIDIA says Blackwell Ultra has 1.5x the dense FP4 Tensor Core FLOPS of Blackwell and 2x the attention-layer acceleration. FP4 matters because inference stacks that run models in 4-bit formats get the most from it. Background in FP4 and NVFP4 vs MXFP4. The FP8 figure on the rack is unchanged, so if you run FP8 or BF16, do not expect a 1.5x gain from compute alone.
Networking
NVIDIA lists 800 Gb/s of network connectivity per GPU on GB300 NVL72, through ConnectX-8 SuperNICs, working with Quantum-X800 InfiniBand or Spectrum-X Ethernet. NVIDIA's DGX GB200 page lists 400 Gb/s ConnectX-7 links, so scale-out bandwidth between racks doubles on paper. This matters when a job spans several racks. Within one rack, NVLink is the fabric and it did not change.
What stayed the same
The rack form factor, the 36 Grace CPUs, the 17 TB of LPDDR5X, the NVLink switch fabric and liquid cooling. NVIDIA does not state a rack power figure on either page. For OEM-listed power on the GB200 rack (Supermicro lists 132 kW), see our GB200 NVL72 guide; do not assume the GB300 figure is identical until your vendor confirms it.
Performance: what is published
Treat every number as the vendor's claim for a workload NVIDIA chose.
- Versus Hopper (NVIDIA's claim). NVIDIA says GB300 NVL72 gives "up to a 50x overall increase in AI factory output performance" against Hopper platforms, built from 10x user responsiveness and 5x throughput per megawatt. The page calls these projections, and the test is DeepSeek-R1 with 32K input and 8K output, GB300 with FP4 and Dynamo disaggregation against H100 with FP8 in-flight batching. The two sides use different software and precision, so this is a system-and-software claim, not a pure chip comparison.
- MLPerf Inference v6.0 (published April 1, 2026). NVIDIA's results post says four GB300 NVL72 systems (288 Blackwell Ultra GPUs) linked with Quantum-X800 InfiniBand reached 2,494,310 tokens per second offline and 1,555,110 tokens per second in the server scenario on DeepSeek-R1.
- Gains over its own earlier results. On DeepSeek-R1 server, per-GPU throughput rose from 2,907 to 8,064 tokens per second per GPU against the previous round (2.77x), and on Llama 3.1 405B server from 170 to 259 (1.52x). NVIDIA attributes the 2.7x to software, and says a cloud partner achieved it. These compare GB300 with its own six-month-old submissions, not with GB200.
Notably, NVIDIA's GB300 page contains no direct GB200 NVL72 comparison, and we found no MLPerf entry that puts the two racks side by side on the same workload. If you see a single "GB300 is N times faster than GB200" number, check its baseline and its source.
Which to choose
Pick GB300 NVL72 when:
- Your model or its KV cache is memory bound, such as long-context serving or very large mixture-of-experts models.
- You run FP4 inference and want the extra dense FP4 compute and attention acceleration.
- You will scale across racks and want the extra network bandwidth per GPU.
Stay with GB200 NVL72 when:
- Your model fits in the GB200 memory and you already have capacity there.
- You depend on FP64 or INT8 throughput, which NVIDIA lists as much higher on GB200.
- You are optimizing for availability rather than the last increment of memory.
If you only need a handful of GPUs rather than a rack, compare the single-GPU parts instead: B300 against B200. For the wider picture, the datacenter GPU guide maps every chip, and the GPU pages are /gpu/nvidia-gb300 and /gpu/nvidia-gb200.
Cost: compare per token, not per hour
Rack pricing depends on your contract and on how fully you load it, and we do not type a rental price here. The live figures are in the box below. For a fair comparison, measure tokens per second on your own model on each platform, then divide the hourly price by tokens per hour. A rack with 1.5x the memory can be cheaper per token even at a higher hourly rate if it lets you drop a sharding stage or serve a longer context, and it can be worse if your model fits comfortably on either.
Rent today
Aquanode manages and optimizes GPUs for training and inference workloads, and you can rent the GPUs below on demand. The box shows live availability; "None right now" means there is no offer at the moment.
What's next
NVIDIA's Vera Rubin NVL72 page lists 72 Rubin GPUs with 288 GB of HBM4 each and 216 TB/s of rack NVLink bandwidth (NVIDIA's Rubin technical blog says 260 TB/s), and says the product is "ramping into full production." Its claims, including one-tenth the cost per million tokens against GB200 NVL72 on Kimi-K2-Thinking, are vendor projections. Read Vera Rubin NVL72 and Rubin vs Blackwell vs Hopper for the detail. The desk-side version of this silicon is covered in DGX Station GB300.
FAQ
What is the difference between GB300 and GB200?
GB300 uses Blackwell Ultra GPUs with about 1.5x the HBM3E and 1.5x the dense FP4 compute, plus 800 Gb/s networking per GPU. The rack layout and 130 TB/s NVLink fabric are the same.
Is GB300 NVL72 faster than GB200 NVL72?
For FP4 inference and memory-bound work, NVIDIA says yes, citing 1.5x dense FP4 and 2x attention acceleration. FP8 and FP16 rack figures are the same on NVIDIA's pages, and we found no published like-for-like MLPerf comparison.
How much memory does GB300 NVL72 have?
NVIDIA lists 20 TB of GPU memory and 17 TB of LPDDR5X CPU memory, for 37 TB of fast memory in total.
Does GB300 NVL72 need liquid cooling?
NVIDIA describes the NVL72 family as liquid-cooled rack-scale systems. NVIDIA's GB300 page does not state a power figure, so check your vendor's facility requirements.
Is the GB300 price higher than GB200?
NVIDIA publishes no list price for either rack. Check the live rental box above for current availability.
Sources
- NVIDIA GB300 NVL72 product page: https://www.nvidia.com/en-us/data-center/gb300-nvl72/
- NVIDIA GB200 NVL72 product page: https://www.nvidia.com/en-us/data-center/gb200-nvl72/
- NVIDIA DGX GB200 product page: https://www.nvidia.com/en-us/data-center/dgx-gb200/
- NVIDIA Blackwell architecture page: https://www.nvidia.com/en-us/data-center/technologies/blackwell-architecture/
- NVIDIA technical blog, MLPerf Inference v6.0 records: https://developer.nvidia.com/blog/nvidia-extreme-co-design-delivers-new-mlperf-inference-records/
- NVIDIA Vera Rubin NVL72 page: https://www.nvidia.com/en-us/data-center/vera-rubin-nvl72/