H200 vs B200 vs GB200: Specs and Benchmarks (2026)

Back
Team Aquanode

Team Aquanode

Sarthak Vaish

Updated OCTOBER 8, 2026Published OCTOBER 8, 2026

B200 is the better choice over H200 when you want more memory per GPU and lower precision (FP4) inference, while GB200 is the same Blackwell silicon wired into a 72-GPU NVLink rack for models that must span many GPUs. NVIDIA's own MLPerf Inference v5.0 write-up shows an eight-GPU B200 system at 3x the Llama 2 70B server throughput of eight H200 GPUs, and GB200 NVL72 at up to 3.4x the per-GPU performance of an eight-GPU H200 system on Llama 3.1 405B.

TL;DR

  • H200 (Hopper): 141 GB HBM3e at 4.8 TB/s, FP8 as the lowest precision. Mature software and widely available.
  • B200 (Blackwell): about 180 GB per GPU in an eight-GPU server (our arithmetic from NVIDIA's DGX B200 page), FP4 support, 1.8 TB/s NVLink.
  • GB200 (Grace Blackwell): Blackwell GPUs paired with Grace CPUs, sold as the 72-GPU NVL72 rack. The difference from B200 is the system, not the GPU generation.
  • Versus H100: NVIDIA claims DGX B200 gives 3x the training and 15x the inference performance of DGX H100 (vendor claim, projected).
  • Verdict: H200 if your model fits and you want proven software. B200 for the best per-node performance. GB200 only when your parallelism needs a rack-wide NVLink domain.

Spec comparison

Figures come from NVIDIA's product pages. Where a number is per GPU but NVIDIA only lists a system total, we divide by the GPU count and mark it as our arithmetic.

SpecH100 SXMH200 SXMB200 (in DGX B200)GB200 (in NVL72)
ArchitectureHopperHopperBlackwellBlackwell + Grace CPU
GPU memory80 GB141 GB HBM3e1,440 GB per 8 GPUs, about 180 GB each (computed)13.4 TB per 72 GPUs, about 186 GB each (computed)
Memory bandwidth3.35 TB/s4.8 TB/s64 TB/s per 8 GPUs, 8 TB/s each (computed)576 TB/s per 72 GPUs, 8 TB/s each (computed)
FP8 Tensor Core3,958 TFLOPS sparse3,958 TFLOPS sparse72 PFLOPS sparse per 8 GPUs, 9 each (computed)720 PFLOPS sparse per 72 GPUs, 10 each (computed)
FP4 Tensor Corenot supportednot supported144 PFLOPS sparse per 8 GPUs, 18 each (computed)1,440 PFLOPS sparse per 72 GPUs, 20 each (computed)
NVLink per GPU900 GB/s900 GB/s1.8 TB/s1.8 TB/s
NVLink domain8 GPUs8 GPUs8 GPUs72 GPUs
TDPup to 700 Wup to 700 Wnot stated on pages we readnot stated on pages we read

Notes on the table:

  • The H100 and H200 FP8 figures are NVIDIA's "with sparsity" numbers. NVIDIA's Blackwell pages state that dense FP8 is half the sparse figure shown, so compare sparse with sparse. The H100 and H200 pages show the same FP8 rating, so the Hopper generation gain from H100 to H200 is memory, not compute.
  • The H200 page is labelled "preliminary specifications, may be subject to change" by NVIDIA.
  • The GB200 per-GPU memory looks slightly higher than B200 because we derive both from system totals; treat 180 GB and 186 GB as approximate, and check the exact SKU with your provider.
  • Hopper has no FP4 hardware support, so a 4-bit model on H200 either falls back to a higher precision or is not usable.

For the head-to-head pages on this site, see B200 vs H200 and B200 vs H100.

What actually differs

Memory

Memory capacity decides whether a model fits without sharding. H200's 141 GB is a 76 percent step up from H100's 80 GB, and NVIDIA says it gives 1.4x the memory bandwidth of H100. B200 goes to about 180 GB per GPU and about 8 TB/s per GPU of bandwidth by our arithmetic, a bit under 1.7 times H200's 4.8 TB/s. In practice that moves the line for a 70B-parameter model at long context or a larger mixture-of-experts model from "two or more GPUs" toward fewer GPUs. See VRAM and HBM for the background.

Precision

Blackwell adds FP4 Tensor Core support, and NVIDIA publishes FP4 figures for B200 and GB200. Hopper tops out at FP8. If your inference stack serves quantized 4-bit models, that is a major reason to move. See FP4, FP8 and quantization, and our post on NVFP4 vs MXFP4. If you serve in FP8 or BF16, the compute gap is smaller than the FP4 headline suggests.

Interconnect

NVLink 4 on Hopper gives 900 GB/s per GPU. NVLink 5 on Blackwell doubles that to 1.8 TB/s. Both B200 servers and H200 servers connect 8 GPUs in one NVLink domain. GB200 NVL72 is the one that extends the domain to 72 GPUs, which matters for large-scale tensor and expert parallelism. See tensor parallelism and what is NVLink.

System shape

B200 is sold mainly as HGX B200 and DGX B200 eight-GPU servers. NVIDIA lists DGX B200 at about 14.3 kW maximum in 10 rack units, air-cooled to a 10 to 35 degrees C operating range. GB200 NVL72 is a liquid-cooled rack that Supermicro lists at 132 kW. That gap is the real cost difference between the two: a B200 server fits a conventional data center, a GB200 rack does not. The HGX vs DGX vs NVL72 post covers how these systems are packaged.

Performance: vendor and MLPerf figures

All performance numbers below are NVIDIA's. Most are NVIDIA-submitted MLPerf results, which are audited by MLCommons, but the workloads are chosen by NVIDIA and per-GPU performance is not a primary MLPerf metric.

  • H200 vs H100 (NVIDIA's claim). On Llama 2 70B, NVIDIA's H200 page claims 1.9x the inference speed of H100 (one GPU each; H100 at batch size 8, H200 at batch size 32, 2K input, 128 output). On GPT-3 175B it claims 1.6x with eight GPUs each. The batch sizes differ, because H200's extra memory lets it run a larger batch, so this is not a same-settings comparison.
  • B200 vs H200 (MLPerf Inference v5.0). In NVIDIA's results post, an eight-GPU Blackwell system delivered 98,443 tokens per second in the Llama 2 70B server scenario against 33,072 for eight H200 GPUs (3x), and 126,845 against 61,802 on Mixtral 8x7B (2.1x). Stable Diffusion XL was 1.6x. On the stricter Llama 2 70B Interactive scenario NVIDIA reports 3.1x.
  • GB200 NVL72 vs H200 (MLPerf Inference v5.0). On Llama 3.1 405B, NVIDIA reports up to 3.4x higher per-GPU performance than an eight-GPU H200 system (2.8x offline, 3.4x server). At system level NVIDIA reports up to 30x, because the rack puts 9x more GPUs in one NVLink domain.
  • B200 vs H100 (NVIDIA's claim). The DGX B200 page claims 3x the training and 15x the inference performance of DGX H100, projected, with an inference test at 50 ms token-to-token latency and 5 s first-token latency, 32,768 input tokens. NVIDIA's inference comparison is per GPU against DGX H100.
  • Cost per token (third-party, via NVIDIA). The DGX B200 page cites SemiAnalysis InferenceX (Q1 2026) at about $0.02 per million tokens on GPT-OSS-120B with TensorRT-LLM, which NVIDIA says is roughly 4.5x cheaper than Hopper at $0.09.

One thing to notice: the MLPerf gap between B200 and H200 (about 2x to 3x on these workloads) is much smaller than the 15x headline for DGX B200 against DGX H100. The headline mixes a generation of hardware with FP4 and software changes and a latency target chosen to favor Blackwell. Expect your own gain to land closer to the audited numbers if you run FP8 or BF16.

Which should you choose

Choose H200 when:

  • Your model fits in 141 GB per GPU and you run FP8 or BF16.
  • You want the widest software and driver maturity and the simplest procurement.
  • Your latency target is relaxed enough that Hopper already meets it.

Choose B200 when:

  • You want more memory per GPU and the best per-node throughput.
  • You serve FP4 models, or plan to.
  • You are training or serving at a scale where eight tightly coupled GPUs is the unit you need. See the B200 guide and B300 vs B200.

Choose GB200 when:

  • One model needs far more than eight GPUs with communication-heavy parallelism, such as large mixture-of-experts serving.
  • Your facility can deliver direct liquid cooling and a rack in the 130 kW class.
  • Read the GB200 NVL72 guide first, and the GB300 comparison if memory is your constraint.

If your model is small, the cheapest GPU that fits it is usually the right answer. The datacenter GPU guide lists every option, and the GPU index shows the live range.

Cost: use throughput, then multiply

We do not type a rental price here. Use cited throughput for a tokens-per-hour figure, then multiply by the live hourly price. For example, NVIDIA's MLPerf v5.0 Llama 2 70B server figure for an eight-GPU B200 system is 98,443 tokens per second, which is about 12,305 tokens per second per GPU on average (our arithmetic), or about 44 million tokens per GPU-hour. The matching H200 figure is 33,072 over eight GPUs, about 4,134 per GPU, or about 14.9 million tokens per GPU-hour. Divide the live hourly price in the box below by each figure to get your cost per million tokens for that benchmark. These are benchmark numbers on one model and will not match your production traffic.

Rent today

Aquanode manages and optimizes GPUs for training and inference workloads, and you can rent the GPUs below on demand. The box shows live availability; a chip that says "None right now" has no offer at the moment.

What's next

NVIDIA's Blackwell Ultra parts (B300 and GB300) raise memory further, covered in our B300 guide. The Rubin generation is next: see Rubin vs Blackwell vs Hopper, which compares all three architectures.

FAQ

Is B200 better than H200?

For raw throughput, yes: NVIDIA's MLPerf v5.0 results show an eight-GPU B200 system at 3x the Llama 2 70B server throughput of eight H200 GPUs. H200 can still be the better value if your model fits and you do not need FP4.

What is the difference between B200 and GB200?

B200 is the Blackwell GPU, usually in eight-GPU servers. GB200 pairs Blackwell GPUs with Grace CPUs, and the NVL72 rack links 72 of them in one NVLink domain.

How much faster is B200 than H100?

NVIDIA claims 3x training and 15x inference for DGX B200 against DGX H100 (vendor projections). The audited MLPerf v5.0 gap against H200 is smaller, so check which benchmark a quoted multiple comes from.

Does H200 support FP4?

No. NVIDIA's H200 page lists FP8 as its lowest Tensor Core precision. FP4 arrives with Blackwell.

How much memory does a B200 have?

NVIDIA lists 1,440 GB across the eight GPUs in DGX B200, about 180 GB per GPU by our arithmetic.

Sources

#datacenter gpu#nvidia blackwell#h200 vs b200#gb200 vs b200#b200 vs h100#gpu comparison

Submit the job. Everything after that is ours.

Sign up in 60 seconds. Pay for the GPU minutes you actually use.

© 2026 Aquanode. All rights reserved.

All trademarks, logos and brand names are the property of their respective owners.