NVIDIA H200 GPU: Specs, VRAM, Price & Benchmarks (2026)
The NVIDIA H200 (141GB HBM3e) launched in 2024. This guide covers its full specs, VRAM, SXM/PCIe differences where they apply, AI performance, and live per-GPU rental pricing on Aquanode.
Short answer on cost: the H200 rents from $3.95 per GPU per hour on Aquanode. Out of capacity right now.
How much VRAM does the H200 have?
The H200 has 141GB HBM3e, with 4.8 TB/s of peak memory bandwidth.
- VRAM: H200 141GB HBM3e vs H100 80GB HBM3
- VRAM: H200 141GB HBM3e vs B200 180GB HBM3e
What fits in 141GB HBM3e of VRAM
| Model | Precision | Fits? |
|---|---|---|
| Llama 3.1 70B | FP16 | ~140GB weights against 141GB. Loads, but leaves almost nothing for KV cache, so serve it on 2 cards or drop to FP8. |
| Llama 3.1 70B | FP8 | ~70GB. Fits with substantial headroom for long-context serving. |
| Mixtral 8x22B | FP16 | ~280GB weights. Still needs 2 cards even at 141GB each. |
| Llama 3.1 405B | INT4 | ~200-230GB. Needs 2 cards; doesn't fit on one H200. |
Approximate, based on published parameter counts and standard bytes-per-parameter rules of thumb (FP16 ≈ 2 bytes/param, INT4 ≈ 0.5-0.6 bytes/param). Real footprint also depends on KV-cache size and framework overhead.
What can the H200 run?
Popular open models from small to frontier scale, with the memory each needs and how many H200 cards (141GB HBM3e each) that takes.
| Model | As published | FP8 | INT4 |
|---|---|---|---|
| Qwen/Qwen3-8B 8.2B | BF16: ~18.3 GB, 1 GPU | FP8: ~9.2 GB, 1 GPU | INT4: ~4.6 GB, 1 GPU |
| Qwen/Qwen2.5-14B-Instruct 14.8B | BF16: ~33 GB, 1 GPU | FP8: ~16.5 GB, 1 GPU | INT4: ~8.3 GB, 1 GPU |
| Qwen/Qwen3-32B 32.8B | BF16: ~73.2 GB, 1 GPU | FP8: ~36.6 GB, 1 GPU | INT4: ~18.3 GB, 1 GPU |
| Qwen/Qwen-72B 72.3B | BF16: ~162 GB, 2 GPUs | FP8: ~80.8 GB, 1 GPU | INT4: ~40.4 GB, 1 GPU |
| MiniMaxAI/MiniMax-M2.7 228.7B | FP8: ~256 GB, 2 GPUs | – | INT4: ~128 GB, 1 GPU |
| deepseek-ai/DeepSeek-R1 684.5B | FP8: ~765 GB, 6 GPUs | – | INT4: ~383 GB, 3 GPUs |
Estimates: weights at the stated precision plus a flat 20% for KV cache and overhead, at a moderate context length. A dash means the precision is not offered for that model (it is already published at that size). INT4 needs a published quantized checkpoint. Open any model for a per-GPU breakdown, or use the H200 VRAM calculator.
H200 VRAM calculator: check which models fit in its memory at each precision.
All models that fit in 141 GB: the open models whose weights and overhead fit, at native, FP8 and INT4 precision.
H200 specs
H200 SXM
| VRAM | 141GB HBM3e |
| Memory bandwidth | 4.8 TB/s |
| TDP | Up to 700W (configurable) |
| Form factor | SXM (DGX H200 / HGX H200 platform only) |
| Interconnect | NVLink 4, 900 GB/s bidirectional |
H200 NVL (PCIe)
| VRAM | 141GB HBM3e |
| Memory bandwidth | 4.8 TB/s |
| TDP | 600W |
| Form factor | PCIe 5.0, air-cooled |
| Interconnect | NVLink bridge, 900 GB/s per GPU (2-4 way) |
Specs sourced from the vendor's public datasheet/product page. See the source.
Related reading: NVIDIA H200 guide, H100 vs H200, H200 vs B200 vs GB200, and H100 and H200: SXM vs NVL vs PCIe.
GPU Glossary: What is VRAM?, HBM, Tensor Cores, CUDA Cores, TFLOPS, NVLink vs PCIe
H200 SXM vs NVL (PCIe): which should you use?
- Memory bandwidth: SXM delivers 4.8 TB/s vs 4.8 TB/s on NVL (PCIe).
- Interconnect: SXM has NVLink 4, 900 GB/s bidirectional. NVL (PCIe) has NVLink bridge, 900 GB/s per GPU (2-4 way).
- Power and form factor: SXM is Up to 700W (configurable) in a SXM (DGX H200 / HGX H200 platform only) form factor. NVL (PCIe) is 600W in a PCIe 5.0, air-cooled form factor.
For large-scale distributed training, the variant with NVLink and the highest memory bandwidth is usually the right choice, since those advantages compound across a multi-GPU cluster. For single-GPU inference or fine-tuning, the lower-power variant often gives the same usable VRAM at a lower hourly rate.
H200 AI performance
Same compute as H100 but 76% more VRAM and 43% more memory bandwidth, so a 70B model fits on one card in FP8 with room for a long KV cache, and it suits longer-context inference where KV-cache size, not compute, is the bottleneck.
- Dense FP16/BF16 tensor throughput: 989 TFLOPS
- Dense FP8 tensor throughput: 1,979 TFLOPS
- Memory bandwidth: 4.8 TB/s
Compute throughput is identical to H100. If a workload is compute-bound rather than memory-bound, the extra VRAM buys nothing and H100 is usually the cheaper hourly rate for the same FLOPS.
See how it stacks up against other cards in the GPU benchmarks and specs table.
The cheapest H200 offer right now ($3.95/GPU/hr) is about 32% below the market median of $5.82/GPU/hr. Out of capacity right now.
H200 price: what does it cost?
Renting. The cheapest current on-demand rate for the H200 on Aquanode is $3.98/GPU/hr. Live rates range from $3.95 to $5.94 per GPU per hour, with a median of $5.82/GPU/hr. Billed by the offer's own terms; the table below shows every live rate. Out of capacity right now.
| Region | $/GPU/hr | Available | VRAM | vCPU | RAM |
|---|---|---|---|---|---|
| – | $3.95 | 0 | 141 GB | – | – |
| Beltsville, Us | $3.98 | 1 | 141 GB | 16 | 180 GB |
| Washington, Us | $4.96 | 1 | 141 GB | 56 | 189 GB |
| Finland | $5.64 | 2 | 141 GB | 44 | 170 GB |
| France | $5.94 | 25 | 141 GB | 16 | 200 GB |
H200 price history
Aquanode stores one snapshot of its GPU price index per UTC day. For the H200 that is 5 days so far, 2026-10-06 to 2026-10-10, so this is a short history, not a long-run trend. The lowest per-GPU rate was $3.95 on both 2026-10-06 and 2026-10-10.
| Day (UTC) | Lowest per GPU hour | Median per GPU hour | Data-center lowest per GPU hour | Offers |
|---|---|---|---|---|
| 2026-10-10 | $3.95 | $5.82 | $3.98 | 68 |
| 2026-10-09 | $3.95 | $5.28 | $3.98 | 52 |
| 2026-10-08 | $3.95 | $5.04 | $3.98 | 58 |
| 2026-10-07 | $3.95 | $5.05 | $3.98 | 53 |
| 2026-10-06 | $3.95 | $5.05 | $3.98 | 58 |
Each row is the stored daily snapshot of the live index, copied as recorded. A dash means that day stored no figure. The same series is available as JSON and summarised in the monthly GPU price report.
How this price is calculated
All prices on this page are normalized to a per-GPU hourly rate using each offer's authoritative GPU count, so that raw price is divided by the number of GPUs it actually covers; some offers report price as already per-GPU, so those are used as-listed. An offer with a missing, zero, or invalid GPU count is excluded entirely rather than published at a guessed rate.
No offers were excluded from this snapshot for a missing or invalid price. No offers were dropped as price outliers in this snapshot.
Only the cheapest qualifying offer per provider is shown in the table above. This page regenerates at most once per hour.
H200 vs H100: how do they compare?
- VRAM: H200 141GB HBM3e vs H100 80GB HBM3
- Memory bandwidth: 4.8 TB/s vs 3.35 TB/s
- Dense FP8 tensor throughput: 1,979 TFLOPS vs 1,979 TFLOPS
On Aquanode right now, H100 starts at $2.19/GPU/hr against H200's $3.95/GPU/hr, about 45% less.
H200 vs B200: how do they compare?
- VRAM: H200 141GB HBM3e vs B200 180GB HBM3e
- Memory bandwidth: 4.8 TB/s vs 8 TB/s
- Dense FP8 tensor throughput: 1,979 TFLOPS vs 4,500 TFLOPS
On Aquanode right now, H200 starts at $3.95/GPU/hr against B200's $8.79/GPU/hr, about 55% less.
Compare the H200 with other GPUs
Compare H200 with
- B200 vs H200
- H100 vs H200
- A100 vs H200
- DGX A100 vs H200
- H200 vs V100
- AMD MI300X vs H200
- H200 vs L40S
- H200 vs L40
- A40 vs H200
- H200 vs RTX A6000
- H200 vs RTX A5000
- H200 vs RTX A4000
- H200 vs L4
- H200 vs T4
- H200 vs RTX 6000 Ada
- H200 vs RTX PRO 6000
- H200 vs RTX PRO 6000 WS
- H200 vs RTX PRO 6000 SE
- H200 vs RTX PRO 5000
- H200 vs RTX 5090
- H200 vs RTX 5080
- H200 vs RTX 5070 Ti
- H200 vs RTX 5070
- H200 vs RTX 5060 Ti
- H200 vs RTX 4090
- H200 vs RTX 4080 Super
- H200 vs RTX 4080
- H200 vs RTX 4070 Ti
- H200 vs RTX 4070 Super
- H200 vs RTX 4070
- H200 vs RTX 4060 Ti
- H200 vs RTX 3090
- H200 vs RTX 3080
- H200 vs RTX 3070
- H200 vs RTX 3060
- B300 vs H200
- AMD MI355X vs H200
- AMD MI325X vs H200
- H100 NVL vs H200
- H200 vs RTX PRO 4500
- H200 vs RTX PRO 4500 SE
- H200 vs RTX PRO 4000
- H200 vs RTX 5880 Ada
- H200 vs RTX 5000 Ada
- H200 vs RTX 4000 Ada
- H200 vs RTX 4000 SFF Ada
- H200 vs RTX 2000 Ada
- H200 vs RTX 4070 Ti Super
- H200 vs RTX A4500
- H200 vs RTX 2080 Ti
- H200 vs RTX 3080 Ti
- H200 vs RTX 3070 Ti
- H200 vs RTX 3060 Ti
- GTX 1080 Ti vs H200
- H200 vs RTX 4060
- H200 vs RTX 5060
- GTX 1660 Super vs H200
- H200 vs RTX 2060
- H200 vs Radeon RX 9070 XT
- H200 vs Radeon RX 6700 XT
- H200 vs RTX 6000
- A16 vs H200
- H200 vs P40
- H200 vs P4
- H200 vs Quadro P2000
- H200 vs Quadro M4000
Get notified when the price drops
GPU supply moves hourly. Tell us what you're waiting for and we'll email you when a matching offer appears across any provider we track.
Rent a H200 on Aquanode
- On-demand instances from $3.95/GPU/hr, billed by the provider's own terms, with no hardware procurement or long-term commitment.
- 73 live offers across 5 regions today.
- Set a price/availability alert above to hear the moment a cheaper or newly-available H200 offer appears.
- Compare every H200 offer side by side, or browse the full multi-provider GPU marketplace.
Good for
Same compute as H100 but 76% more VRAM and 43% more memory bandwidth, so a 70B model fits on one card in FP8 with room for a long KV cache, and it suits longer-context inference where KV-cache size, not compute, is the bottleneck.
Not good for
Compute throughput is identical to H100. If a workload is compute-bound rather than memory-bound, the extra VRAM buys nothing and H100 is usually the cheaper hourly rate for the same FLOPS.
H200 FAQs
How much VRAM does the H200 have?
The H200 has 141GB HBM3e, with 4.8 TB/s of peak memory bandwidth.
What is the NVIDIA H200?
The NVIDIA H200 is a GPU released in 2024, with 141GB HBM3e of memory and a Up to 700W (configurable) power envelope. See the full spec table above for interconnect, form factor and tensor-throughput details.
How much does it cost to rent a H200?
Live H200 rental prices currently range from $3.95 to $5.94 per GPU per hour, with a median of $5.82 per GPU per hour. Out of capacity right now.
What's the cheapest H200 rate?
The lowest current H200 rate on Aquanode is $3.95 per GPU per hour in . Out of capacity right now.
What is the difference between H200 SXM and NVL (PCIe)?
SXM has 4.8 TB/s of memory bandwidth and NVLink 4, 900 GB/s bidirectional. NVL (PCIe) has 4.8 TB/s and NVLink bridge, 900 GB/s per GPU (2-4 way). Both carry 141GB HBM3e.
How does the H200 compare to the H100?
VRAM: H200 141GB HBM3e vs H100 80GB HBM3 Memory bandwidth: 4.8 TB/s vs 3.35 TB/s Dense FP8 tensor throughput: 1,979 TFLOPS vs 1,979 TFLOPS On Aquanode right now, H100 starts at $2.19/GPU/hr against H200's $3.95/GPU/hr, about 45% less. See the full H200 vs H100 comparison for a shared-provider price breakdown.
How does the H200 compare to the B200?
VRAM: H200 141GB HBM3e vs B200 180GB HBM3e Memory bandwidth: 4.8 TB/s vs 8 TB/s Dense FP8 tensor throughput: 1,979 TFLOPS vs 4,500 TFLOPS On Aquanode right now, H200 starts at $3.95/GPU/hr against B200's $8.79/GPU/hr, about 55% less. See the full H200 vs B200 comparison for a shared-provider price breakdown.
What is the H200 good for?
Same compute as H100 but 76% more VRAM and 43% more memory bandwidth, so a 70B model fits on one card in FP8 with room for a long KV cache, and it suits longer-context inference where KV-cache size, not compute, is the bottleneck.
What are the H200's limitations?
Compute throughput is identical to H100. If a workload is compute-bound rather than memory-bound, the extra VRAM buys nothing and H100 is usually the cheaper hourly rate for the same FLOPS.
Is renting cheaper than buying?
Renting avoids the upfront hardware cost and lets you match spend to actual usage. A rented H200 at $3.95/hr only costs money while it's running, whereas buying ties up capital in hardware that keeps depreciating whether it's in use or not. Which is cheaper depends on how continuously you'd run it; short or bursty workloads usually favor renting. Out of capacity right now.
How is the H200 price calculated?
All prices are normalized to a per-GPU hourly rate using each offer's authoritative GPU count, which the raw price is divided by; some offers report price as already per-GPU. Offers whose price can't be safely normalized, or whose rate is an extreme outlier against the rest of the market, are excluded.
H200 price by region
Related guides
Other models in the same generation, then the rest of the GPU index.