The NVIDIA B300 is the Blackwell Ultra GPU: 288 GB of HBM3e, 8 TB/s of memory bandwidth and 15 PFLOPS of dense NVFP4 compute per GPU, with a maximum power of up to 1,400 W. It is the same two-die Blackwell design as the B200 with more memory, 1.5x the dense FP4 throughput and faster attention, according to NVIDIA, and it is sold as the HGX B300 baseboard, the DGX B300 system and, with a Grace CPU, as the GB300 superchip.
This guide covers:
- B300 specs, memory and power, with a source for every number
- What changed from B200, and what did not
- DGX B300 and HGX B300 explained
- Published MLPerf and vendor figures, labelled as such
- Infrastructure needs and when B300 is the right pick
TL;DR
- Memory: 288 GB HBM3e per GPU and 8 TB/s bandwidth, which NVIDIA says is 3.6x the memory of an H100.
- Compute: 15 PFLOPS dense NVFP4 per GPU against 10 PFLOPS for Blackwell. FP8 dense stays at 5 PFLOPS, the same as Blackwell.
- What got cut: NVIDIA's HGX table shows FP64 falling to 10 TFLOPS per eight-GPU board (B200: 296) and INT8 to 3 POPS (B200: 72). B300 is built for low-precision AI, not scientific FP64.
- Power: up to 1,400 W per GPU, and about 14 kW for a DGX B300 system.
- Verdict: pick B300 for long-context and large-model inference where memory capacity and FP4 throughput set your cost per token. Stay on B200 if your model fits in its memory and you use FP8 or BF16.
Part of our datacenter GPU guide.
B300 specs at a glance
Per-GPU figures below come from NVIDIA's technical blog on Blackwell Ultra. Board and system figures come from the NVIDIA HGX and DGX B300 pages.
| Spec | B200 (Blackwell) | B300 (Blackwell Ultra) |
|---|---|---|
| Transistors | 208 billion | 208 billion |
| Process | TSMC 4NP | TSMC 4NP |
| Memory per GPU | 180 GB HBM3e (HGX and DGX config) | 288 GB HBM3e |
| Memory bandwidth per GPU | Up to 8 TB/s | 8 TB/s |
| NVFP4 dense per GPU | 10 PFLOPS | 15 PFLOPS |
| NVFP4 sparse per GPU | 20 PFLOPS | 20 PFLOPS |
| FP8 dense per GPU | 5 PFLOPS | 5 PFLOPS |
| NVLink per GPU | 1.8 TB/s | 1.8 TB/s |
| Max power per GPU | Configurable up to 1,000 W | Up to 1,400 W |
| 8-GPU board FP4 (sparse / dense) | 144 / 72 PFLOPS | 144 / 108 PFLOPS |
| 8-GPU board total memory | 1.4 TB | 2.1 TB |
| 8-GPU board network bandwidth | 0.8 TB/s | 1.6 TB/s |
Three things in this table need a note. First, the per-GPU rows come from NVIDIA's blog and the eight-GPU rows from its HGX page, so they do not divide exactly: 72 dense PFLOPS over eight B200 GPUs is 9, not the blog's 10, and NVIDIA does not explain the gap.
Second, 288 GB per chip versus 2.1 TB per board. Eight times 288 GB is about 2.3 TB, but NVIDIA's HGX B300 and DGX B300 pages both list 2.1 TB total. NVIDIA's pages do not explain the difference. Plan around the 2.1 TB board figure when sizing an eight-GPU node, and check the exact usable memory with nvidia-smi on the machine you get.
Third, dense versus sparse FP4. The headline "144 PFLOPS" on both DGX pages is the sparse number. The dense number for B300 is 108 PFLOPS per eight-GPU board, and for B200 it is 72, a ratio of 1.5, which matches NVIDIA's claim that DGX B300 raises dense FP4 by 1.5x. Sparse FP4 is identical on the two chips. Unless your weights use structured sparsity, dense is the number to compare.
What Blackwell Ultra changes
NVIDIA's Blackwell Ultra technical blog (published August 22, 2025, updated September 24, 2025) lists these changes and constants:
- Same silicon basics. Two reticle-sized dies connected by NV-HBI at 10 TB/s, presented as one CUDA device, with up to 160 SMs and 640 fifth-generation tensor cores. SM count varies by SKU.
- More memory. 288 GB of HBM3e in 12-high stacks, against 80 GB on H100. The blog's memory bullet and its update note disagree on the stack count (eight versus twelve), so we do not quote a stack count here.
- More NVFP4 throughput. 15 PFLOPS dense, which the blog calls 1.5x Blackwell and 7.5x Hopper.
- Faster attention. The blog doubles special function unit throughput for the exponential operation used in softmax, from 5 to 10.7 tera-exponentials per second, and claims up to 2x faster attention-layer compute than Blackwell. Attention is the part of a transformer that grows with context length, so this matters most for long prompts and reasoning models.
- Same NVLink. NVLink 5 at 1.8 TB/s per GPU, with a fabric that scales to 576 GPUs.
- Faster networking. The Grace Blackwell Ultra superchip uses ConnectX-8 at 800 Gb/s. NVIDIA's DGX B300 page lists eight ConnectX-8 adapters at up to 800 Gb/s each, double the 400 Gb/s ConnectX-7 on DGX B200.
- PCIe Gen 6 x16 at 256 GB/s bidirectional.
If NVFP4 is new to you, read our NVFP4 vs MXFP4 explainer and the FP4 glossary entry first. FP4 is why B300's headline numbers are so large.
DGX B300 and HGX B300
Like B200, B300 comes in several shapes:
HGX B300 is the eight-GPU baseboard that server makers build into their own systems. NVIDIA lists it as shipping now, with 2.1 TB of total memory, 14.4 TB/s of total NVLink bandwidth and 1.6 TB/s of network bandwidth.
DGX B300 is NVIDIA's finished system. Per NVIDIA's page it has eight Blackwell Ultra SXM GPUs, Intel Xeon 6776P processors, two BlueField-3 DPUs at up to 400 Gb/s, 8 x 3.84 TB of NVMe storage, and a 10 RU chassis that fits NVIDIA MGX racks and traditional enterprise racks. Power is about 14 kW, with AC and DC options.
GB300 (Grace Blackwell Ultra) pairs the GPU with a Grace CPU over NVLink-C2C at 900 GB/s, and the GB300 NVL72 rack links 72 GPUs. That is a different purchase and a different cooling plan. We compare it with its predecessor in GB300 NVL72 vs GB200 NVL72, and the form factors in HGX vs DGX vs NVL72.
A B300 node you rent by the hour is almost always the HGX or DGX eight-GPU shape.
Performance: what is published
Every figure here is a vendor claim or an MLPerf submission. We have not benchmarked B300 ourselves.
MLPerf Inference v5.1 (September 9, 2025). On DeepSeek-R1, NVIDIA reported these per-GPU figures, computed by dividing system throughput by accelerator count (per-GPU is not a primary MLPerf metric):
| System | Offline (tokens/s/GPU) | Server (tokens/s/GPU) |
|---|---|---|
| DGX H200, 8 GPUs (Hopper) | 1,253 | 556 |
| GB200 NVL72 (Blackwell) | 4,024 | 2,327 |
| GB300 NVL72 (Blackwell Ultra) | 5,842 | 2,907 |
NVIDIA's reading: GB300 NVL72 is about 45 percent higher offline and about 25 percent higher in the server scenario than GB200 NVL72. The Hopper results were not verified by MLCommons. Note these are rack-scale GB300 numbers, which use the same Blackwell Ultra GPU as B300 but a different system, so they are not what an eight-GPU HGX B300 will give you. NVIDIA says most DeepSeek-R1 weights were quantized to NVFP4 and the KV cache to FP8 in these runs.
MLPerf Inference v6.0 (April 2026). NVIDIA's DGX B300 page states Blackwell Ultra systems delivered 2.5 million tokens per second on DeepSeek-R1, up to 2.7x higher than the Blackwell Ultra debut submissions six months earlier. StorageReview's coverage identifies that run as 288 Blackwell Ultra GPUs across four GB300 NVL72 racks. The takeaway is that software maturity moved Blackwell Ultra a long way inside six months on the same silicon.
Cost per token and efficiency, as NVIDIA states them. The DGX B300 page cites SemiAnalysis InferenceX (Q1 2026): up to 50x higher throughput per megawatt and up to 35x lower cost per token than Hopper for low-latency agentic workloads, and $0.24 per million tokens at 102 tokens per second per user on DeepSeek-R1, using NVIDIA Dynamo, TensorRT-LLM and multi-token prediction. These are vendor-presented, workload-specific and measured on the best software NVIDIA can assemble.
Generation claim. NVIDIA says DGX B300 boosts dense FP4 performance by 1.5x and attention performance by 2x over DGX B200.
Infrastructure needs
- Power. NVIDIA's blog gives a maximum of up to 1,400 W per GPU for Blackwell Ultra (Hopper: 700 W, Blackwell: 1,200 W). The DGX B300 system is about 14 kW. Verify the exact figure for any HGX B300 server, since the vendor chooses the configuration.
- Cooling and racks. DGX B300 is a 10 RU chassis for MGX or enterprise racks. Rack-scale GB300 is a separate, denser design.
- Networking. ConnectX-8 at up to 800 Gb/s per GPU is the point of the upgrade, and it only pays off with a fabric that can use it. See NVLink vs InfiniBand and RDMA.
- Software. You need a stack built for Blackwell Ultra: a recent CUDA, NCCL (glossary), and an inference engine with NVFP4 kernels, such as TensorRT-LLM. NVIDIA names Dynamo and TensorRT-LLM in its B300 cost figures.
When to choose B300
B300 fits when:
- Your model or its KV cache does not fit comfortably in 180 GB. 288 GB per GPU means fewer GPUs per replica for very large mixture-of-experts models and long contexts.
- You run reasoning or long-context inference, where the 2x attention claim matters.
- You have, or are building, an NVFP4 inference pipeline and want the 1.5x dense FP4 headroom. See quantization for the accuracy side.
B200 fits when:
- Your workload is FP8 or BF16. NVIDIA's table shows no dense FP8, FP16 or BF16 gain for B300 over B200.
- Your model fits in 180 GB and you are not memory or attention bound.
- You need FP64 or INT8 throughput, which B300 trades away.
The full decision framework is in B300 vs B200. You can also read the spec-level B300 page and the B200 page.
Cost: how to think about it
Start from a published throughput, then apply the live price. The cleanest anchor is NVIDIA's MLPerf v5.1 DeepSeek-R1 result for GB300 NVL72: 5,842 tokens per second per GPU offline and 2,907 in the server scenario. Multiplied by 3,600 seconds that is about 21.0 million and 10.5 million tokens per GPU-hour (computed), on a rack-scale system with benchmark-tuned software.
Treat those as ceilings. Your batch size, latency target and engine version will land you below them. Take your measured tokens per GPU-hour, divide the live hourly price in the box below by it, and you have a cost per million tokens. Run the same sum for B200. B300's extra memory can also lower cost indirectly, by letting one GPU hold a model replica that would need two B200s. Current prices are on /pricing and the GPU index.
Rent today
Aquanode manages and optimizes GPUs for training and inference workloads, and you can rent the GPUs in the box below on demand. The box shows live availability, including when a chip has no offer right now.
What's next
The next generation after Blackwell Ultra is Rubin. NVIDIA's Rubin page describes the Vera Rubin platform as in full production, and its MLPerf page claims Vera Rubin NVL72 delivers up to 3.7x higher throughput than GB300 NVL72 in MLPerf Inference v6.1. Those are vendor figures for rack-scale systems, and NVIDIA publishes no availability window on the page we read. Our Rubin guide separates what is announced from what you can use today. Until then, B300 and B200 are the Blackwell options.
FAQ
How much memory does the B300 have?
288 GB of HBM3e per GPU according to NVIDIA's Blackwell Ultra blog, with 8 TB/s of bandwidth. NVIDIA's HGX B300 and DGX B300 pages list 2.1 TB for an eight-GPU board.
Is B300 the same as Blackwell Ultra?
Yes. B300 is the Blackwell Ultra data center GPU. GB300 is the Grace CPU plus Blackwell Ultra superchip, and GB300 NVL72 is the 72-GPU rack.
How much faster is B300 than B200?
NVIDIA states 1.5x dense FP4 and 2x attention performance for DGX B300 over DGX B200. Dense FP8, FP16 and BF16 figures on NVIDIA's HGX page are the same for both. Real speedups depend on whether your workload is FP4, memory-bound or attention-bound.
How much power does a B300 draw?
NVIDIA's blog lists up to 1,400 W per GPU, and its DGX B300 page lists about 14 kW for the system.
Does B300 replace B200?
Not for every workload. B300 trades FP64 and INT8 throughput for low-precision AI performance, per NVIDIA's HGX table, so B200 remains the better-balanced part for some jobs. See B300 vs B200.
What is DGX B300?
NVIDIA's complete eight-GPU Blackwell Ultra system, with Intel Xeon 6776P CPUs, ConnectX-8 networking, BlueField-3 DPUs and NVIDIA's software, in a 10 RU chassis.
Sources
- NVIDIA DGX B300: https://www.nvidia.com/en-us/data-center/dgx-b300/
- NVIDIA HGX platform: https://www.nvidia.com/en-us/data-center/hgx/
- NVIDIA, Inside NVIDIA Blackwell Ultra: https://developer.nvidia.com/blog/inside-nvidia-blackwell-ultra-the-chip-powering-the-ai-factory-era/
- NVIDIA, Blackwell Ultra sets new inference records in MLPerf debut (MLPerf Inference v5.1): https://developer.nvidia.com/blog/nvidia-blackwell-ultra-sets-new-inference-records-in-mlperf-debut/
- NVIDIA MLPerf benchmarks page (MLPerf Inference v6.1 claims): https://www.nvidia.com/en-us/data-center/resources/mlperf-benchmarks/
- NVIDIA HGX AI Factory reference architecture (B200 memory and bandwidth): https://docs.nvidia.com/enterprise-reference-architectures/hgx-ai-factory-h100-h200-b200/latest/components.html
- NVIDIA DGX B200 (B200 comparison): https://www.nvidia.com/en-us/data-center/dgx-b200/
- NVIDIA Rubin platform page: https://www.nvidia.com/en-us/data-center/technologies/rubin/
- MLPerf Inference v6.0 coverage (288 GPU, four GB300 NVL72 racks): https://www.storagereview.com/news/nvidia-sets-mlperf-inference-v6-0-records-with-blackwell-ultra-platform