NVIDIA B200 Guide: Specs, VRAM, DGX B200 (2026)

Back
Team Aquanode

Team Aquanode

Sarthak Vaish

Updated OCTOBER 8, 2026Published OCTOBER 8, 2026

The NVIDIA B200 is the first Blackwell data center GPU: two reticle-sized dies in one package, 180 GB of HBM3e per GPU in NVIDIA's HGX and DGX configurations, up to 8 TB/s of memory bandwidth, and fifth-generation NVLink at 1.8 TB/s per GPU. It is sold mainly as an eight-GPU board (HGX B200) or as the NVIDIA DGX B200 system, as of October 2026.

This guide covers:

  • B200 specs, VRAM and bandwidth, with a source for every number
  • What DGX B200 and HGX B200 are, and how they differ
  • Vendor-published performance figures, labelled as vendor claims
  • Power, cooling and networking a B200 node needs
  • When B200 is the right pick over H200 or B300

TL;DR

  • Memory: 180 GB HBM3e per GPU and 1.44 TB per eight-GPU node, with bandwidth of up to 8 TB/s per GPU, per NVIDIA's HGX reference architecture.
  • Compute: NVIDIA lists 144 PFLOPS of FP4 (sparse) and 72 PFLOPS dense for the eight-GPU DGX B200, with FP8 at 72 PFLOPS sparse.
  • Interconnect: NVLink 5 at 1.8 TB/s GPU to GPU, 14.4 TB/s total across the baseboard.
  • Power: configurable up to 1,000 W per GPU, and about 14.3 kW for a full DGX B200 system.
  • Verdict: B200 is the default Blackwell choice for large-model training and high-throughput inference when you need more memory and FP4 than Hopper but do not need the 288 GB of B300.

This post is part of our datacenter GPU guide, which maps every accelerator generation we cover.

B200 specs at a glance

The table compares B200 with the two Hopper parts it replaces. B200 and H100/H200 figures come from NVIDIA's HGX reference architecture page. The FP4 row is the baseboard figure NVIDIA publishes for the whole eight-GPU board, because NVIDIA's HGX page gives compute at board level.

SpecH100 SXMH200 SXMB200 SXM
ArchitectureHopperHopperBlackwell
Memory per GPU80 GB HBM3141 GB HBM3e180 GB HBM3e
Memory per 8-GPU node640 GB1.1 TB1.44 TB
Memory bandwidth per GPU3.35 TB/s4.8 TB/sUp to 8 TB/s
NVLink generation4th4th5th
GPU-to-GPU NVLink900 GB/s900 GB/s1,800 GB/s
Baseboard NVLink total7.2 TB/s7.2 TB/s14.4 TB/s
Max power per GPUnot listednot listedConfigurable up to 1,000 W
FP4 tensor, 8-GPU board (sparse / dense)not supportednot supported144 / 72 PFLOPS

Two notes on reading this table. First, "up to 8 TB/s" is NVIDIA's own wording for B200 bandwidth; some third-party sheets print 7.7 TB/s, which we could not find in an NVIDIA document, so we use NVIDIA's number. Second, the 180 GB figure is the HGX and DGX B200 configuration. NVIDIA's DGX B200 page lists 1,440 GB total across eight GPUs, which is 180 GB each.

If you want the VRAM story in more depth, our glossary entry on VRAM explains what actually has to fit: weights, KV cache and activations. For B200, 180 GB means a 70B-parameter model in FP8 (about 70 GB of weights) leaves well over 100 GB for KV cache on a single GPU.

What is inside a B200

NVIDIA's Blackwell architecture page says Blackwell GPUs pack 208 billion transistors, built from two reticle-limited dies joined by a 10 TB/s chip-to-chip link. The two dies are presented to software as one GPU, so you program a B200 the way you program any single CUDA device.

Beyond raw size, NVIDIA calls out several architecture features on the same page:

  • Second-generation Transformer Engine. It adds support for 4-bit floating point (FP4) alongside FP8, using what NVIDIA calls microscaling formats. See the FP4 glossary entry and our comparison of NVFP4 and MXFP4.
  • Fifth-generation NVLink. NVIDIA says it scales to 576 GPUs, and it doubles per-GPU bandwidth over Hopper's 900 GB/s.
  • A dedicated RAS engine. Reliability, availability and serviceability hardware that flags likely faults early to cut downtime.
  • Confidential computing. NVIDIA describes Blackwell as the first TEE-I/O-capable GPU, with throughput nearly identical to unencrypted modes.
  • A decompression engine for database workloads (LZ4, Snappy, Deflate), connected to a Grace CPU at 900 GB/s in the superchip designs.

The tensor cores are the part that matters most for AI work: they are where FP4, FP8 and BF16 math actually run.

DGX B200, HGX B200 and the rack-scale options

"B200" gets used for three different things, and they are not interchangeable.

HGX B200 is the eight-GPU baseboard that server makers build into their own servers. NVIDIA's HGX page lists eight Blackwell SXM GPUs, 1.4 TB of total memory, 144 PFLOPS of FP4 sparse (72 dense) and 14.4 TB/s of total NVLink bandwidth.

DGX B200 is NVIDIA's own finished system built on that baseboard. Per NVIDIA's DGX B200 page it adds two Intel Xeon Platinum 8570 processors (112 cores total), 2 TB of system memory (configurable to 4 TB), eight ConnectX-7 network adapters at up to 400 Gb/s each, two BlueField-3 DPUs, and a 10 RU chassis. It also ships with NVIDIA AI Enterprise and Mission Control software.

GB200 NVL72 is a different product: a rack-scale system that links 72 Blackwell GPUs into one NVLink domain (NVIDIA lists 130 TB/s of rack NVLink bandwidth). It uses Blackwell GPUs but is not a B200 server. For that path, read our GB200 NVL72 guide, and for how the three form factors compare, see HGX vs DGX vs NVL72.

For most teams renting capacity, the HGX or DGX B200 shape is what a "B200 node" means: eight GPUs, one machine.

Performance: what NVIDIA publishes

Everything below is a vendor claim or an MLPerf result, and each one is linked in Sources. We do not add our own benchmarks.

Generation-over-generation, as NVIDIA states it. NVIDIA's DGX B200 page says the system delivers 3x the training performance and 15x the inference performance of DGX H100. Both figures are footnoted as projected performance subject to change. The inference footnote assumes 50 ms token-to-token latency, a 5 s first-token latency, 32,768 input tokens and 1,028 output tokens, comparing per-GPU performance. The training footnote compares 4,096-node clusters at 32,768-GPU scale. Treat these as best-case, workload-specific claims, not a promise for your model.

Cost per token. The same page cites SemiAnalysis InferenceX results (April 2026): roughly $0.02 per million tokens at 55 tokens per second per user on a Blackwell system running TensorRT-LLM, against $0.09 for Hopper on vLLM, which NVIDIA calls roughly 4.5x cheaper. The page ties the $0.02 figure to the GPT-OSS-120B model. This is a vendor-presented third-party benchmark on one model; your numbers depend on your model, batch size and latency target.

MLPerf Inference. In MLPerf Inference v5.1 (published September 9, 2025), NVIDIA reported that disaggregated serving on GB200 NVL72 gave nearly 1.5x per-GPU throughput over aggregated serving on DGX B200 for Llama 3.1 405B. That is a statement about serving architecture, and a reminder that the same GPU can deliver very different throughput depending on how the software is arranged. NVIDIA also said it used NVFP4 extensively across its Blackwell DeepSeek-R1 and Llama submissions.

What B200 has that H100 does not. Based on NVIDIA's own table: 2.25x the memory (180 GB against 80 GB), up to 2.4x the bandwidth, 2x the per-GPU NVLink bandwidth, and a native 4-bit tensor path that Hopper lacks. Those are spec ratios, not application speedups.

Infrastructure needs

A B200 node is a data center machine, not a workstation part.

  • Power. NVIDIA configures B200 up to 1,000 W per GPU, and lists about 14.3 kW maximum system power for DGX B200. Plan for a rack with high-density power, not a standard office circuit.
  • Cooling. DGX B200 is a 10 RU air-cooled system with an operating temperature range of 10 to 35 degrees C (50 to 90 F) per NVIDIA. The rack-scale NVL72 line is a separate product.
  • Networking. DGX B200 uses eight ConnectX-7 adapters at up to 400 Gb/s for compute traffic, plus BlueField-3 DPUs. Multi-node training depends on that fabric; see NVLink vs InfiniBand and RDMA for why.
  • Software. CUDA, a recent TensorRT-LLM, vLLM or SGLang build with Blackwell kernels, and NCCL (glossary) for multi-GPU communication. Older framework versions may not use FP4 at all, so check your stack before assuming you get the FP4 numbers.

When to choose B200

Choose B200 when:

  • You are training or fine-tuning large models and want 180 GB per GPU plus NVLink 5 bandwidth across an eight-GPU node. For parallelism strategies that stress the interconnect, see tensor parallelism.
  • You serve large mixture-of-experts models, where weights are large and memory capacity drives how many GPUs you need.
  • You want to use FP4 inference on a mature Blackwell part. See quantization for the accuracy trade-offs.

Look at H200 instead when your model fits in 141 GB, you are not using FP4, and you want the older, widely supported software path. Our H200 vs B200 vs GB200 comparison walks through the split, and there is a spec-by-spec B200 vs H200 page.

Look at B300 instead when you need more memory per GPU or the extra dense FP4 throughput. NVIDIA's DGX B300 page says B300 raises dense FP4 performance by 1.5x and attention performance by 2x over DGX B200. We cover the trade-offs in B300 vs B200.

Cost: how to think about it

Hourly price is only half of a cost model. The other half is throughput. A published figure you can anchor on is NVIDIA's MLPerf Inference v5.1 DeepSeek-R1 result for GB200 NVL72, which NVIDIA reports as 4,024 offline tokens per second per GPU (computed by dividing system throughput by 72 GPUs; per-GPU numbers are not a primary MLPerf metric). At 3,600 seconds an hour that is about 14.5 million tokens per GPU-hour (computed), on a rack-scale GB200 system with software tuned for the benchmark.

Your own throughput will almost certainly be lower, because benchmark runs use tuned batch sizes and offline scenarios. Use the figure to sanity-check the order of magnitude, then multiply tokens per GPU-hour by the live hourly price in the box below to get a cost per million tokens. Compare that against the same calculation for H200 and H100 before you commit to a generation. You can also see all current prices on /pricing and the GPU index.

Rent today

Aquanode manages and optimizes GPUs for training and inference workloads, and you can rent the GPUs in the box below on demand. Prices are live, so the box shows what is available right now.

See the full B200 page for specs and current availability.

What's next

B200 has a direct successor in the same Blackwell family: B300 (Blackwell Ultra), which NVIDIA lists with 288 GB of HBM3e per GPU and 15 PFLOPS of dense FP4. After that comes Rubin. NVIDIA's Rubin page describes the Vera Rubin platform as in full production, but it publishes no availability window and no quantified comparison to Blackwell. See our Rubin guide for what is and is not confirmed.

FAQ

How much VRAM does the B200 have?

180 GB of HBM3e per GPU in NVIDIA's HGX and DGX B200 configurations, which is 1.44 TB across an eight-GPU node. NVIDIA's DGX B200 page lists the node total as 1,440 GB.

What is the difference between DGX B200 and HGX B200?

HGX B200 is the eight-GPU baseboard that server vendors build into their own machines. DGX B200 is NVIDIA's complete system built on that baseboard, with CPUs, networking, storage and NVIDIA software included.

What is B200 memory bandwidth?

NVIDIA's HGX reference architecture lists up to 8 TB/s per GPU and up to 64 TB/s across an eight-GPU node. NVIDIA's DGX B200 page also gives 64 TB/s for the system.

Does B200 support FP4?

Yes. Blackwell's second-generation Transformer Engine supports 4-bit floating point. NVIDIA lists 144 PFLOPS of FP4 for DGX B200, which is the sparse figure; the dense figure is 72 PFLOPS. Always check whether a quoted number is sparse or dense. The FP4 glossary explains the format.

How much power does a B200 use?

NVIDIA documents a configurable maximum of 1,000 W per B200 GPU and about 14.3 kW for a complete DGX B200 system.

Is B200 better than H100?

On specs, yes: more memory, more bandwidth, faster NVLink and a 4-bit tensor path. NVIDIA claims 3x training and 15x inference performance over DGX H100 for projected, footnoted workloads. For a model that fits an H100 and does not use FP4, the gap will be much smaller, so measure your own workload. There is a side-by-side B200 vs H100 page.

Sources

#datacenter gpu#nvidia blackwell#b200#dgx b200#hgx b200#blackwell gpu specs

Submit the job. Everything after that is ours.

Sign up in 60 seconds. Pay for the GPU minutes you actually use.

© 2026 Aquanode. All rights reserved.

All trademarks, logos and brand names are the property of their respective owners.