H200 GPU rental price
NVIDIA H200 · 141 GB VRAM. Live per-GPU pricing aggregated across 5 providers and 5 regions.
The H200 (141GB HBM3e) currently rents from $3.59 per GPU per hour on Aquanode, across 5 providers.
The cheapest H200 offer right now (RunPod, $3.59/GPU/hr) is about 22% below the market median of $4.59/GPU/hr.
H200 specs
H200 full specs
| Architecture | NVIDIA, launched 2024 |
| VRAM | 141GB HBM3e |
| Memory bandwidth | 4.8 TB/s |
| FP16 / BF16 tensor throughput | 989 TFLOPS (peak, dense) |
| FP8 tensor throughput | 1,979 TFLOPS (peak, dense) |
| Interconnect | NVLink 4, 900 GB/s bidirectional |
| TDP | Up to 700W (configurable) |
| Form factor | SXM, PCIe (H200 NVL) |
Specs sourced from the vendor's public datasheet. See the source.
What fits in 141GB HBM3e of VRAM
| Model | Precision | Fits? |
|---|---|---|
| Llama 3.1 70B | FP16 | ~140GB weights. Fits on a single card with room for a real KV cache. |
| Llama 3.1 70B | FP8 | ~70GB. Fits with substantial headroom for long-context serving. |
| Mixtral 8x22B | FP16 | ~280GB weights. Still needs 2 cards even at 141GB each. |
| Llama 3.1 405B | INT4 | ~200-230GB. Needs 2 cards; doesn't fit on one H200. |
Approximate, based on published parameter counts and standard bytes-per-parameter rules of thumb (FP16 ≈ 2 bytes/param, INT4 ≈ 0.5-0.6 bytes/param). Real footprint also depends on KV-cache size and framework overhead.
Good for
Same compute as H100 but 76% more VRAM and 43% more memory bandwidth, so it's the better single-card fit for a 70B model in FP16 or a longer-context inference workload where KV-cache size, not compute, is the bottleneck.
Not good for
Compute throughput is identical to H100. If a workload is compute-bound rather than memory-bound, the extra VRAM buys nothing and H100 is usually the cheaper hourly rate for the same FLOPS.
Get notified when the price drops
GPU supply moves hourly. Tell us what you're waiting for and we'll email you when a matching offer appears across any provider we track.
How this price is calculated
All prices on this page are normalized to a per-GPU hourly rate using each offer's authoritative GPU count, so that raw price is divided by the number of GPUs it actually covers; Akash reports its price as already per-GPU, so it is used as-listed. An offer with a missing, zero, or invalid GPU count is excluded entirely rather than published at a guessed rate.
No offers were excluded from this snapshot for a missing or invalid price. No offers were dropped as price outliers in this snapshot.
Only the cheapest qualifying offer per provider is shown in the table above. This page regenerates at most once per hour.
Frequently asked questions
How much does it cost to rent a H200?
Live H200 rental prices currently range from $3.59 to $4.59 per GPU per hour across 5 providers, with a median of $4.59 per GPU per hour.
Which provider has the cheapest H200?
RunPod currently offers the lowest H200 rate on Aquanode's marketplace at $3.59 per GPU per hour in Unknown.
How is the H200 price calculated?
All prices are normalized to a per-GPU hourly rate using each offer's authoritative GPU count, which the raw price is divided by; Akash reports its price as already per-GPU. Offers whose price can't be safely normalized, or whose rate is an extreme outlier against the rest of the market, are excluded.
How much does an H200 cost per hour?
Live H200 rental rates on Aquanode currently range from $3.59 to $4.59 per GPU per hour, with a median of $4.59/hr. See the live table above for current per-provider pricing.
How much VRAM does an H200 have?
The H200 has 141GB HBM3e, with 4.8 TB/s of peak memory bandwidth.
Is renting cheaper than buying?
Renting avoids the upfront hardware cost and lets you match spend to actual usage. A rented H200 at $3.59/hr only costs money while it's running, whereas buying ties up capital in hardware that keeps depreciating whether it's in use or not. Which is cheaper depends on how continuously you'd run it; short or bursty workloads usually favor renting.