GB300 NVL72 vs GB200 NVL72: Specs and Differences (2026)

Back
Team Aquanode

Team Aquanode

Sarthak Vaish

Updated OCTOBER 8, 2026Published OCTOBER 8, 2026

GB300 NVL72 is the Blackwell Ultra upgrade of GB200 NVL72: the same 72-GPU, 36-CPU liquid-cooled rack layout, but with about 1.5 times the HBM3E (20 TB against 13.4 TB), 1.5 times the dense FP4 compute, and 800 Gb/s of network bandwidth per GPU. NVLink bandwidth stays at 130 TB/s, so the rack-scale fabric did not change; the GPUs and the memory around them did.

TL;DR

  • Same rack, better GPUs. Both racks hold 72 GPUs and 36 Grace CPUs and share 130 TB/s of NVLink bandwidth. GB300 swaps in Blackwell Ultra GPUs.
  • Memory is the headline. NVIDIA lists 20 TB of GPU memory for GB300 NVL72 against 13.4 TB for GB200 NVL72, which NVIDIA describes as 1.5x larger HBM3E.
  • Compute gain is FP4 and attention. NVIDIA claims 1.5x more dense FP4 Tensor Core FLOPS and 2x attention-layer acceleration. FP8 and FP16 rack figures are unchanged on NVIDIA's pages.
  • Some precision was traded away. NVIDIA's GB300 page lists much lower INT8 and FP64 numbers than the GB200 page does.
  • Verdict: pick GB300 for long-context and large-model inference where memory limits you. GB200 is fine if you already have it or the memory fits.

Spec comparison

All figures below are from NVIDIA's product pages for each rack. FP4 is the sparse figure unless marked.

SpecGB200 NVL72GB300 NVL72
GPUs72 Blackwell72 Blackwell Ultra
CPUs36 Grace (2,592 Arm Neoverse V2 cores)36 Grace (2,592 Arm Neoverse V2 cores)
GPU memory13.4 TB HBM3E20 TB
GPU memory bandwidth576 TB/sup to 576 TB/s
CPU memory17 TB LPDDR5X, 14 TB/s17 TB LPDDR5X, 14 TB/s
Total fast memory30 TB37 TB
NVLink bandwidth130 TB/s130 TB/s
FP4 Tensor Core1,440 PFLOPS sparse, 720 dense1,440 PFLOPS sparse, 1,080 dense
FP8/FP6 Tensor Core720 PFLOPS720 PFLOPS
FP16/BF16 Tensor Core360 PFLOPS360 PFLOPS
INT8 Tensor Core720 POPS24 POPS
FP642,880 TFLOPS100 TFLOPS
Network per GPU400 Gb/s ConnectX-7 (DGX GB200 page)800 Gb/s ConnectX-8
Rack powernot stated by NVIDIAnot stated by NVIDIA

Three notes on reading that table:

  1. NVIDIA's GB300 NVL72 page lists FP4 as 1,440 and 1,080 PFLOPS, with a footnote marking the second figure "without sparsity"; the page does not label the two columns more directly. The 1,080 dense figure is also exactly 1.5 times the 720 dense figure on the GB200 page. The 30 TB and 37 TB totals for fast memory (GPU plus CPU memory) come from NVIDIA's DGX GB200 page and the GB300 page respectively.
  2. Dividing by 72 GPUs (our arithmetic) gives about 186 GB of HBM3E per GPU on GB200 and about 278 GB on GB300.
  3. The INT8 and FP64 drop is real on NVIDIA's pages. Blackwell Ultra is tuned for low-precision AI inference. If your workload leans on FP64 (some HPC and simulation codes), check the numbers before assuming GB300 is a straight upgrade.

What changed and why it matters

Memory per GPU

The 1.5x increase in HBM3E is the change most buyers will feel. A bigger HBM stack per GPU means larger models, longer contexts and bigger KV caches fit before you have to shard further. See HBM and KV cache in the glossary. The same step is what separates the single-GPU B300 from the B200; our B300 vs B200 post covers it per GPU.

FP4 and attention

NVIDIA says Blackwell Ultra has 1.5x the dense FP4 Tensor Core FLOPS of Blackwell and 2x the attention-layer acceleration. FP4 matters because inference stacks that run models in 4-bit formats get the most from it. Background in FP4 and NVFP4 vs MXFP4. The FP8 figure on the rack is unchanged, so if you run FP8 or BF16, do not expect a 1.5x gain from compute alone.

Networking

NVIDIA lists 800 Gb/s of network connectivity per GPU on GB300 NVL72, through ConnectX-8 SuperNICs, working with Quantum-X800 InfiniBand or Spectrum-X Ethernet. NVIDIA's DGX GB200 page lists 400 Gb/s ConnectX-7 links, so scale-out bandwidth between racks doubles on paper. This matters when a job spans several racks. Within one rack, NVLink is the fabric and it did not change.

What stayed the same

The rack form factor, the 36 Grace CPUs, the 17 TB of LPDDR5X, the NVLink switch fabric and liquid cooling. NVIDIA does not state a rack power figure on either page. For OEM-listed power on the GB200 rack (Supermicro lists 132 kW), see our GB200 NVL72 guide; do not assume the GB300 figure is identical until your vendor confirms it.

Performance: what is published

Treat every number as the vendor's claim for a workload NVIDIA chose.

  • Versus Hopper (NVIDIA's claim). NVIDIA says GB300 NVL72 gives "up to a 50x overall increase in AI factory output performance" against Hopper platforms, built from 10x user responsiveness and 5x throughput per megawatt. The page calls these projections, and the test is DeepSeek-R1 with 32K input and 8K output, GB300 with FP4 and Dynamo disaggregation against H100 with FP8 in-flight batching. The two sides use different software and precision, so this is a system-and-software claim, not a pure chip comparison.
  • MLPerf Inference v6.0 (published April 1, 2026). NVIDIA's results post says four GB300 NVL72 systems (288 Blackwell Ultra GPUs) linked with Quantum-X800 InfiniBand reached 2,494,310 tokens per second offline and 1,555,110 tokens per second in the server scenario on DeepSeek-R1.
  • Gains over its own earlier results. On DeepSeek-R1 server, per-GPU throughput rose from 2,907 to 8,064 tokens per second per GPU against the previous round (2.77x), and on Llama 3.1 405B server from 170 to 259 (1.52x). NVIDIA attributes the 2.7x to software, and says a cloud partner achieved it. These compare GB300 with its own six-month-old submissions, not with GB200.

Notably, NVIDIA's GB300 page contains no direct GB200 NVL72 comparison, and we found no MLPerf entry that puts the two racks side by side on the same workload. If you see a single "GB300 is N times faster than GB200" number, check its baseline and its source.

Which to choose

Pick GB300 NVL72 when:

  • Your model or its KV cache is memory bound, such as long-context serving or very large mixture-of-experts models.
  • You run FP4 inference and want the extra dense FP4 compute and attention acceleration.
  • You will scale across racks and want the extra network bandwidth per GPU.

Stay with GB200 NVL72 when:

  • Your model fits in the GB200 memory and you already have capacity there.
  • You depend on FP64 or INT8 throughput, which NVIDIA lists as much higher on GB200.
  • You are optimizing for availability rather than the last increment of memory.

If you only need a handful of GPUs rather than a rack, compare the single-GPU parts instead: B300 against B200. For the wider picture, the datacenter GPU guide maps every chip, and the GPU pages are /gpu/nvidia-gb300 and /gpu/nvidia-gb200.

Cost: compare per token, not per hour

Rack pricing depends on your contract and on how fully you load it, and we do not type a rental price here. The live figures are in the box below. For a fair comparison, measure tokens per second on your own model on each platform, then divide the hourly price by tokens per hour. A rack with 1.5x the memory can be cheaper per token even at a higher hourly rate if it lets you drop a sharding stage or serve a longer context, and it can be worse if your model fits comfortably on either.

Rent today

Aquanode manages and optimizes GPUs for training and inference workloads, and you can rent the GPUs below on demand. The box shows live availability; "None right now" means there is no offer at the moment.

What's next

NVIDIA's Vera Rubin NVL72 page lists 72 Rubin GPUs with 288 GB of HBM4 each and 216 TB/s of rack NVLink bandwidth (NVIDIA's Rubin technical blog says 260 TB/s), and says the product is "ramping into full production." Its claims, including one-tenth the cost per million tokens against GB200 NVL72 on Kimi-K2-Thinking, are vendor projections. Read Vera Rubin NVL72 and Rubin vs Blackwell vs Hopper for the detail. The desk-side version of this silicon is covered in DGX Station GB300.

FAQ

What is the difference between GB300 and GB200?

GB300 uses Blackwell Ultra GPUs with about 1.5x the HBM3E and 1.5x the dense FP4 compute, plus 800 Gb/s networking per GPU. The rack layout and 130 TB/s NVLink fabric are the same.

Is GB300 NVL72 faster than GB200 NVL72?

For FP4 inference and memory-bound work, NVIDIA says yes, citing 1.5x dense FP4 and 2x attention acceleration. FP8 and FP16 rack figures are the same on NVIDIA's pages, and we found no published like-for-like MLPerf comparison.

How much memory does GB300 NVL72 have?

NVIDIA lists 20 TB of GPU memory and 17 TB of LPDDR5X CPU memory, for 37 TB of fast memory in total.

Does GB300 NVL72 need liquid cooling?

NVIDIA describes the NVL72 family as liquid-cooled rack-scale systems. NVIDIA's GB300 page does not state a power figure, so check your vendor's facility requirements.

Is the GB300 price higher than GB200?

NVIDIA publishes no list price for either rack. Check the live rental box above for current availability.

Sources

#datacenter gpu#nvidia blackwell#gb300#gb200#nvl72#blackwell ultra

Submit the job. Everything after that is ours.

Sign up in 60 seconds. Pay for the GPU minutes you actually use.

© 2026 Aquanode. All rights reserved.

All trademarks, logos and brand names are the property of their respective owners.