NVIDIA B200 GPU: Specs, VRAM, Price & Benchmarks (2026)

The NVIDIA B200 (180GB HBM3e) launched in 2024. This guide covers its full specs, VRAM, SXM/PCIe differences where they apply, AI performance, and live per-GPU rental pricing on Aquanode.

Short answer on cost: the B200 rents from $8.79 per GPU per hour on Aquanode. Out of capacity right now.

How much VRAM does the B200 have?

The B200 has 180GB HBM3e, with 8 TB/s of peak memory bandwidth.

  • VRAM: B200 180GB HBM3e vs H200 141GB HBM3e
  • VRAM: B200 180GB HBM3e vs AMD MI300X 192GB HBM3

What fits in 180GB HBM3e of VRAM

ModelPrecisionFits?
Llama 3.1 70BFP16~140GB. Fits on one card with headroom to spare.
Llama 3.1 405BFP8~410GB. Needs 3+ cards even at 180GB each.
Llama 3.1 405BINT4~200-230GB. Still needs 2 cards.
Mixtral 8x22BFP8~140GB. Fits on a single card.

Approximate, based on published parameter counts and standard bytes-per-parameter rules of thumb (FP16 ≈ 2 bytes/param, INT4 ≈ 0.5-0.6 bytes/param). Real footprint also depends on KV-cache size and framework overhead.

What can the B200 run?

Popular open models from small to frontier scale, with the memory each needs and how many B200 cards (180GB HBM3e each) that takes.

ModelAs publishedFP8INT4
Qwen/Qwen3-8B 8.2BBF16: ~18.3 GB, 1 GPUFP8: ~9.2 GB, 1 GPUINT4: ~4.6 GB, 1 GPU
Qwen/Qwen2.5-14B-Instruct 14.8BBF16: ~33 GB, 1 GPUFP8: ~16.5 GB, 1 GPUINT4: ~8.3 GB, 1 GPU
Qwen/Qwen3-32B 32.8BBF16: ~73.2 GB, 1 GPUFP8: ~36.6 GB, 1 GPUINT4: ~18.3 GB, 1 GPU
Qwen/Qwen-72B 72.3BBF16: ~162 GB, 1 GPUFP8: ~80.8 GB, 1 GPUINT4: ~40.4 GB, 1 GPU
MiniMaxAI/MiniMax-M2.7 228.7BFP8: ~256 GB, 2 GPUs–INT4: ~128 GB, 1 GPU
deepseek-ai/DeepSeek-R1 684.5BFP8: ~765 GB, 5 GPUs–INT4: ~383 GB, 3 GPUs

Estimates: weights at the stated precision plus a flat 20% for KV cache and overhead, at a moderate context length. A dash means the precision is not offered for that model (it is already published at that size). INT4 needs a published quantized checkpoint. Open any model for a per-GPU breakdown, or use the B200 VRAM calculator.

B200 VRAM calculator: check which models fit in its memory at each precision.

All models that fit in 141 GB: the open models whose weights and overhead fit, at native, FP8 and INT4 precision.

$8.79/GPU/hrOut of capacity right now
Lowest / GPU / hr
$8.79/GPU/hr
Median / GPU / hr
$9.92/GPU/hr
p90 / GPU / hr
12
Live offers
Last updated: 2026-10-10 21:05:44 UTCRefreshes hourly0 offer(s) excluded from this snapshot

B200 specs

ArchitectureNVIDIA, launched 2024
VRAM180GB HBM3e
Memory bandwidth8 TB/s
FP16 / BF16 tensor throughput2,250 TFLOPS (peak, dense)
FP8 tensor throughput4,500 TFLOPS (peak, dense)
InterconnectNVLink 5, 1.8 TB/s per GPU
TDPUp to 1,000W
Form factorSXM

Specs sourced from the vendor's public datasheet/product page. See the source.

Related reading: NVIDIA B200 guide, B300 vs B200, H200 vs B200 vs GB200, and The best GPUs for AI, ranked.

GPU Glossary: What is VRAM?, HBM, Tensor Cores, CUDA Cores, TFLOPS, NVLink vs PCIe

B200 AI performance

The highest-throughput card on this marketplace: more than double H100's FP8 throughput and over double the memory bandwidth, plus native FP4 tensor-core support for the newest FP4-quantized model formats. Built for frontier-scale training and the highest-throughput inference deployments.

  • Dense FP16/BF16 tensor throughput: 2,250 TFLOPS
  • Dense FP8 tensor throughput: 4,500 TFLOPS
  • Memory bandwidth: 8 TB/s

Supply is thin and the hourly rate is the highest on the marketplace. It's the wrong card for a workload that fits comfortably on an H100 or H200, since the extra throughput goes unused and the cost premium doesn't pay for itself.

See how it stacks up against other cards in the GPU benchmarks and specs table.

B200 price: what does it cost?

Renting. The cheapest current on-demand rate for the B200 on Aquanode is $8.79/GPU/hr. Live rates range from $8.79 to $9.92 per GPU per hour, with a median of $8.79/GPU/hr. Billed by the offer's own terms; the table below shows every live rate. Out of capacity right now.

Region$/GPU/hrAvailableVRAMvCPURAM
–$8.790180 GB––
United States$9.3523180 GB20224 GB
Maryland, Us$9.982180 GB48283.5 GB

B200 price history

Aquanode stores one snapshot of its GPU price index per UTC day. For the B200 that is 5 days so far, 2026-10-06 to 2026-10-10, so this is a short history, not a long-run trend. The lowest per-GPU rate went from $7.47 on 2026-10-06 to $8.16 on 2026-10-10.

Day (UTC)Lowest per GPU hourMedian per GPU hourData-center lowest per GPU hourOffers
2026-10-10$8.16$8.79$8.1615
2026-10-09$8.08$8.79$8.0817
2026-10-08$7.47$7.47$7.4713
2026-10-07$7.47$7.47$7.4712
2026-10-06$7.47$7.47$7.4710

Each row is the stored daily snapshot of the live index, copied as recorded. A dash means that day stored no figure. The same series is available as JSON and summarised in the monthly GPU price report.

How this price is calculated

All prices on this page are normalized to a per-GPU hourly rate using each offer's authoritative GPU count, so that raw price is divided by the number of GPUs it actually covers; some offers report price as already per-GPU, so those are used as-listed. An offer with a missing, zero, or invalid GPU count is excluded entirely rather than published at a guessed rate.

No offers were excluded from this snapshot for a missing or invalid price. No offers were dropped as price outliers in this snapshot.

Only the cheapest qualifying offer per provider is shown in the table above. This page regenerates at most once per hour.

B200 vs H200: how do they compare?

  • VRAM: B200 180GB HBM3e vs H200 141GB HBM3e
  • Memory bandwidth: 8 TB/s vs 4.8 TB/s
  • Dense FP8 tensor throughput: 4,500 TFLOPS vs 1,979 TFLOPS

On Aquanode right now, H200 starts at $3.95/GPU/hr against B200's $8.79/GPU/hr, about 55% less.

Full B200 vs H200 price comparison

B200 vs AMD MI300X: how do they compare?

  • VRAM: B200 180GB HBM3e vs AMD MI300X 192GB HBM3
  • Memory bandwidth: 8 TB/s vs 5.3 TB/s
  • Dense FP8 tensor throughput: 4,500 TFLOPS vs 2,614.9 TFLOPS

On Aquanode right now, AMD MI300X starts at $2.63/GPU/hr against B200's $8.79/GPU/hr, about 70% less.

Full B200 vs AMD MI300X price comparison

Compare the B200 with other GPUs

Compare B200 with

Get notified when the price drops

GPU supply moves hourly. Tell us what you're waiting for and we'll email you when a matching offer appears across any provider we track.

One email per matching alert. Unsubscribe any time.

Rent a B200 on Aquanode

  • On-demand instances from $8.79/GPU/hr, billed by the provider's own terms, with no hardware procurement or long-term commitment.
  • 12 live offers across 3 regions today.
  • Set a price/availability alert above to hear the moment a cheaper or newly-available B200 offer appears.
  • Compare every B200 offer side by side, or browse the full multi-provider GPU marketplace.

Good for

The highest-throughput card on this marketplace: more than double H100's FP8 throughput and over double the memory bandwidth, plus native FP4 tensor-core support for the newest FP4-quantized model formats. Built for frontier-scale training and the highest-throughput inference deployments.

Not good for

Supply is thin and the hourly rate is the highest on the marketplace. It's the wrong card for a workload that fits comfortably on an H100 or H200, since the extra throughput goes unused and the cost premium doesn't pay for itself.

B200 FAQs

How much VRAM does the B200 have?

The B200 has 180GB HBM3e, with 8 TB/s of peak memory bandwidth.

What is the NVIDIA B200?

The NVIDIA B200 is a GPU released in 2024, with 180GB HBM3e of memory and a Up to 1,000W power envelope. See the full spec table above for interconnect, form factor and tensor-throughput details.

How much does it cost to rent a B200?

Live B200 rental prices currently range from $8.79 to $9.92 per GPU per hour, with a median of $8.79 per GPU per hour. Out of capacity right now.

What's the cheapest B200 rate?

The lowest current B200 rate on Aquanode is $8.79 per GPU per hour in . Out of capacity right now.

How does the B200 compare to the H200?

VRAM: B200 180GB HBM3e vs H200 141GB HBM3e Memory bandwidth: 8 TB/s vs 4.8 TB/s Dense FP8 tensor throughput: 4,500 TFLOPS vs 1,979 TFLOPS On Aquanode right now, H200 starts at $3.95/GPU/hr against B200's $8.79/GPU/hr, about 55% less. See the full B200 vs H200 comparison for a shared-provider price breakdown.

How does the B200 compare to the AMD MI300X?

VRAM: B200 180GB HBM3e vs AMD MI300X 192GB HBM3 Memory bandwidth: 8 TB/s vs 5.3 TB/s Dense FP8 tensor throughput: 4,500 TFLOPS vs 2,614.9 TFLOPS On Aquanode right now, AMD MI300X starts at $2.63/GPU/hr against B200's $8.79/GPU/hr, about 70% less. See the full B200 vs AMD MI300X comparison for a shared-provider price breakdown.

What is the B200 good for?

The highest-throughput card on this marketplace: more than double H100's FP8 throughput and over double the memory bandwidth, plus native FP4 tensor-core support for the newest FP4-quantized model formats. Built for frontier-scale training and the highest-throughput inference deployments.

What are the B200's limitations?

Supply is thin and the hourly rate is the highest on the marketplace. It's the wrong card for a workload that fits comfortably on an H100 or H200, since the extra throughput goes unused and the cost premium doesn't pay for itself.

Is renting cheaper than buying?

Renting avoids the upfront hardware cost and lets you match spend to actual usage. A rented B200 at $8.79/hr only costs money while it's running, whereas buying ties up capital in hardware that keeps depreciating whether it's in use or not. Which is cheaper depends on how continuously you'd run it; short or bursty workloads usually favor renting. Out of capacity right now.

How is the B200 price calculated?

All prices are normalized to a per-GPU hourly rate using each offer's authoritative GPU count, which the raw price is divided by; some offers report price as already per-GPU. Offers whose price can't be safely normalized, or whose rate is an extreme outlier against the rest of the market, are excluded.

B200 price by region

Related guides

Other models in the same generation, then the rest of the GPU index.

Submit the job. Everything after that is ours.

Sign up in 60 seconds. Pay for the GPU minutes you actually use.

© 2026 Aquanode. All rights reserved.

All trademarks, logos and brand names are the property of their respective owners.