NVIDIA B200 GPU: Specs, VRAM, Price & Benchmarks (2026)
The NVIDIA B200 (180GB HBM3e) launched in 2024. This guide covers its full specs, VRAM, SXM/PCIe differences where they apply, AI performance, and live per-GPU rental pricing on Aquanode.
Short answer on cost: the B200 rents from $8.79 per GPU per hour on Aquanode. Out of capacity right now.
How much VRAM does the B200 have?
The B200 has 180GB HBM3e, with 8 TB/s of peak memory bandwidth.
- VRAM: B200 180GB HBM3e vs H200 141GB HBM3e
- VRAM: B200 180GB HBM3e vs AMD MI300X 192GB HBM3
What fits in 180GB HBM3e of VRAM
| Model | Precision | Fits? |
|---|---|---|
| Llama 3.1 70B | FP16 | ~140GB. Fits on one card with headroom to spare. |
| Llama 3.1 405B | FP8 | ~410GB. Needs 3+ cards even at 180GB each. |
| Llama 3.1 405B | INT4 | ~200-230GB. Still needs 2 cards. |
| Mixtral 8x22B | FP8 | ~140GB. Fits on a single card. |
Approximate, based on published parameter counts and standard bytes-per-parameter rules of thumb (FP16 ≈ 2 bytes/param, INT4 ≈ 0.5-0.6 bytes/param). Real footprint also depends on KV-cache size and framework overhead.
What can the B200 run?
Popular open models from small to frontier scale, with the memory each needs and how many B200 cards (180GB HBM3e each) that takes.
| Model | As published | FP8 | INT4 |
|---|---|---|---|
| Qwen/Qwen3-8B 8.2B | BF16: ~18.3 GB, 1 GPU | FP8: ~9.2 GB, 1 GPU | INT4: ~4.6 GB, 1 GPU |
| Qwen/Qwen2.5-14B-Instruct 14.8B | BF16: ~33 GB, 1 GPU | FP8: ~16.5 GB, 1 GPU | INT4: ~8.3 GB, 1 GPU |
| Qwen/Qwen3-32B 32.8B | BF16: ~73.2 GB, 1 GPU | FP8: ~36.6 GB, 1 GPU | INT4: ~18.3 GB, 1 GPU |
| Qwen/Qwen-72B 72.3B | BF16: ~162 GB, 1 GPU | FP8: ~80.8 GB, 1 GPU | INT4: ~40.4 GB, 1 GPU |
| MiniMaxAI/MiniMax-M2.7 228.7B | FP8: ~256 GB, 2 GPUs | – | INT4: ~128 GB, 1 GPU |
| deepseek-ai/DeepSeek-R1 684.5B | FP8: ~765 GB, 5 GPUs | – | INT4: ~383 GB, 3 GPUs |
Estimates: weights at the stated precision plus a flat 20% for KV cache and overhead, at a moderate context length. A dash means the precision is not offered for that model (it is already published at that size). INT4 needs a published quantized checkpoint. Open any model for a per-GPU breakdown, or use the B200 VRAM calculator.
B200 VRAM calculator: check which models fit in its memory at each precision.
All models that fit in 141 GB: the open models whose weights and overhead fit, at native, FP8 and INT4 precision.
B200 specs
| Architecture | NVIDIA, launched 2024 |
| VRAM | 180GB HBM3e |
| Memory bandwidth | 8 TB/s |
| FP16 / BF16 tensor throughput | 2,250 TFLOPS (peak, dense) |
| FP8 tensor throughput | 4,500 TFLOPS (peak, dense) |
| Interconnect | NVLink 5, 1.8 TB/s per GPU |
| TDP | Up to 1,000W |
| Form factor | SXM |
Specs sourced from the vendor's public datasheet/product page. See the source.
Related reading: NVIDIA B200 guide, B300 vs B200, H200 vs B200 vs GB200, and The best GPUs for AI, ranked.
GPU Glossary: What is VRAM?, HBM, Tensor Cores, CUDA Cores, TFLOPS, NVLink vs PCIe
B200 AI performance
The highest-throughput card on this marketplace: more than double H100's FP8 throughput and over double the memory bandwidth, plus native FP4 tensor-core support for the newest FP4-quantized model formats. Built for frontier-scale training and the highest-throughput inference deployments.
- Dense FP16/BF16 tensor throughput: 2,250 TFLOPS
- Dense FP8 tensor throughput: 4,500 TFLOPS
- Memory bandwidth: 8 TB/s
Supply is thin and the hourly rate is the highest on the marketplace. It's the wrong card for a workload that fits comfortably on an H100 or H200, since the extra throughput goes unused and the cost premium doesn't pay for itself.
See how it stacks up against other cards in the GPU benchmarks and specs table.
B200 price: what does it cost?
Renting. The cheapest current on-demand rate for the B200 on Aquanode is $8.79/GPU/hr. Live rates range from $8.79 to $9.92 per GPU per hour, with a median of $8.79/GPU/hr. Billed by the offer's own terms; the table below shows every live rate. Out of capacity right now.
| Region | $/GPU/hr | Available | VRAM | vCPU | RAM |
|---|---|---|---|---|---|
| – | $8.79 | 0 | 180 GB | – | – |
| United States | $9.35 | 23 | 180 GB | 20 | 224 GB |
| Maryland, Us | $9.98 | 2 | 180 GB | 48 | 283.5 GB |
B200 price history
Aquanode stores one snapshot of its GPU price index per UTC day. For the B200 that is 5 days so far, 2026-10-06 to 2026-10-10, so this is a short history, not a long-run trend. The lowest per-GPU rate went from $7.47 on 2026-10-06 to $8.16 on 2026-10-10.
| Day (UTC) | Lowest per GPU hour | Median per GPU hour | Data-center lowest per GPU hour | Offers |
|---|---|---|---|---|
| 2026-10-10 | $8.16 | $8.79 | $8.16 | 15 |
| 2026-10-09 | $8.08 | $8.79 | $8.08 | 17 |
| 2026-10-08 | $7.47 | $7.47 | $7.47 | 13 |
| 2026-10-07 | $7.47 | $7.47 | $7.47 | 12 |
| 2026-10-06 | $7.47 | $7.47 | $7.47 | 10 |
Each row is the stored daily snapshot of the live index, copied as recorded. A dash means that day stored no figure. The same series is available as JSON and summarised in the monthly GPU price report.
How this price is calculated
All prices on this page are normalized to a per-GPU hourly rate using each offer's authoritative GPU count, so that raw price is divided by the number of GPUs it actually covers; some offers report price as already per-GPU, so those are used as-listed. An offer with a missing, zero, or invalid GPU count is excluded entirely rather than published at a guessed rate.
No offers were excluded from this snapshot for a missing or invalid price. No offers were dropped as price outliers in this snapshot.
Only the cheapest qualifying offer per provider is shown in the table above. This page regenerates at most once per hour.
B200 vs H200: how do they compare?
- VRAM: B200 180GB HBM3e vs H200 141GB HBM3e
- Memory bandwidth: 8 TB/s vs 4.8 TB/s
- Dense FP8 tensor throughput: 4,500 TFLOPS vs 1,979 TFLOPS
On Aquanode right now, H200 starts at $3.95/GPU/hr against B200's $8.79/GPU/hr, about 55% less.
B200 vs AMD MI300X: how do they compare?
- VRAM: B200 180GB HBM3e vs AMD MI300X 192GB HBM3
- Memory bandwidth: 8 TB/s vs 5.3 TB/s
- Dense FP8 tensor throughput: 4,500 TFLOPS vs 2,614.9 TFLOPS
On Aquanode right now, AMD MI300X starts at $2.63/GPU/hr against B200's $8.79/GPU/hr, about 70% less.
Compare the B200 with other GPUs
Compare B200 with
- B200 vs H200
- B200 vs H100
- A100 vs B200
- B200 vs DGX A100
- B200 vs V100
- AMD MI300X vs B200
- B200 vs L40S
- B200 vs L40
- A40 vs B200
- B200 vs RTX A6000
- B200 vs RTX A5000
- B200 vs RTX A4000
- B200 vs L4
- B200 vs T4
- B200 vs RTX 6000 Ada
- B200 vs RTX PRO 6000
- B200 vs RTX PRO 6000 WS
- B200 vs RTX PRO 6000 SE
- B200 vs RTX PRO 5000
- B200 vs RTX 5090
- B200 vs RTX 5080
- B200 vs RTX 5070 Ti
- B200 vs RTX 5070
- B200 vs RTX 5060 Ti
- B200 vs RTX 4090
- B200 vs RTX 4080 Super
- B200 vs RTX 4080
- B200 vs RTX 4070 Ti
- B200 vs RTX 4070 Super
- B200 vs RTX 4070
- B200 vs RTX 4060 Ti
- B200 vs RTX 3090
- B200 vs RTX 3080
- B200 vs RTX 3070
- B200 vs RTX 3060
- B200 vs B300
- AMD MI355X vs B200
- AMD MI325X vs B200
- B200 vs H100 NVL
- B200 vs RTX PRO 4500
- B200 vs RTX PRO 4500 SE
- B200 vs RTX PRO 4000
- B200 vs RTX 5880 Ada
- B200 vs RTX 5000 Ada
- B200 vs RTX 4000 Ada
- B200 vs RTX 4000 SFF Ada
- B200 vs RTX 2000 Ada
- B200 vs RTX 4070 Ti Super
- B200 vs RTX A4500
- B200 vs RTX 2080 Ti
- B200 vs RTX 3080 Ti
- B200 vs RTX 3070 Ti
- B200 vs RTX 3060 Ti
- B200 vs GTX 1080 Ti
- B200 vs RTX 4060
- B200 vs RTX 5060
- B200 vs GTX 1660 Super
- B200 vs RTX 2060
- B200 vs Radeon RX 9070 XT
- B200 vs Radeon RX 6700 XT
- B200 vs RTX 6000
- A16 vs B200
- B200 vs P40
- B200 vs P4
- B200 vs Quadro P2000
- B200 vs Quadro M4000
Get notified when the price drops
GPU supply moves hourly. Tell us what you're waiting for and we'll email you when a matching offer appears across any provider we track.
Rent a B200 on Aquanode
- On-demand instances from $8.79/GPU/hr, billed by the provider's own terms, with no hardware procurement or long-term commitment.
- 12 live offers across 3 regions today.
- Set a price/availability alert above to hear the moment a cheaper or newly-available B200 offer appears.
- Compare every B200 offer side by side, or browse the full multi-provider GPU marketplace.
Good for
The highest-throughput card on this marketplace: more than double H100's FP8 throughput and over double the memory bandwidth, plus native FP4 tensor-core support for the newest FP4-quantized model formats. Built for frontier-scale training and the highest-throughput inference deployments.
Not good for
Supply is thin and the hourly rate is the highest on the marketplace. It's the wrong card for a workload that fits comfortably on an H100 or H200, since the extra throughput goes unused and the cost premium doesn't pay for itself.
B200 FAQs
How much VRAM does the B200 have?
The B200 has 180GB HBM3e, with 8 TB/s of peak memory bandwidth.
What is the NVIDIA B200?
The NVIDIA B200 is a GPU released in 2024, with 180GB HBM3e of memory and a Up to 1,000W power envelope. See the full spec table above for interconnect, form factor and tensor-throughput details.
How much does it cost to rent a B200?
Live B200 rental prices currently range from $8.79 to $9.92 per GPU per hour, with a median of $8.79 per GPU per hour. Out of capacity right now.
What's the cheapest B200 rate?
The lowest current B200 rate on Aquanode is $8.79 per GPU per hour in . Out of capacity right now.
How does the B200 compare to the H200?
VRAM: B200 180GB HBM3e vs H200 141GB HBM3e Memory bandwidth: 8 TB/s vs 4.8 TB/s Dense FP8 tensor throughput: 4,500 TFLOPS vs 1,979 TFLOPS On Aquanode right now, H200 starts at $3.95/GPU/hr against B200's $8.79/GPU/hr, about 55% less. See the full B200 vs H200 comparison for a shared-provider price breakdown.
How does the B200 compare to the AMD MI300X?
VRAM: B200 180GB HBM3e vs AMD MI300X 192GB HBM3 Memory bandwidth: 8 TB/s vs 5.3 TB/s Dense FP8 tensor throughput: 4,500 TFLOPS vs 2,614.9 TFLOPS On Aquanode right now, AMD MI300X starts at $2.63/GPU/hr against B200's $8.79/GPU/hr, about 70% less. See the full B200 vs AMD MI300X comparison for a shared-provider price breakdown.
What is the B200 good for?
The highest-throughput card on this marketplace: more than double H100's FP8 throughput and over double the memory bandwidth, plus native FP4 tensor-core support for the newest FP4-quantized model formats. Built for frontier-scale training and the highest-throughput inference deployments.
What are the B200's limitations?
Supply is thin and the hourly rate is the highest on the marketplace. It's the wrong card for a workload that fits comfortably on an H100 or H200, since the extra throughput goes unused and the cost premium doesn't pay for itself.
Is renting cheaper than buying?
Renting avoids the upfront hardware cost and lets you match spend to actual usage. A rented B200 at $8.79/hr only costs money while it's running, whereas buying ties up capital in hardware that keeps depreciating whether it's in use or not. Which is cheaper depends on how continuously you'd run it; short or bursty workloads usually favor renting. Out of capacity right now.
How is the B200 price calculated?
All prices are normalized to a per-GPU hourly rate using each offer's authoritative GPU count, which the raw price is divided by; some offers report price as already per-GPU. Offers whose price can't be safely normalized, or whose rate is an extreme outlier against the rest of the market, are excluded.
B200 price by region
Related guides
Other models in the same generation, then the rest of the GPU index.