Cloud GPU Price Trends 2026: Why Rates Are Rising Again

Back
Team Aquanode

Team Aquanode

Sarthak Vaish

AUGUST 23, 2026

Last updated: August 23, 2026. This section gets a fresh pass roughly every quarter, right after NVIDIA's earnings call and whenever a new pricing snapshot changes the picture below. If you're reading this more than three months after the date above, treat the direction as directionally right and the numbers as stale — check the live marketplace for current rates.

If you've been renting GPUs since 2023, you've watched prices fall for two straight years and assumed that trend just continues. It doesn't. The reference index most people building GPU infrastructure now watch — SemiAnalysis's H100 rental composite — bottomed around $2.79 to $2.80 per GPU-hour in mid-to-late 2025, then reversed.

TL;DR: Cloud GPU prices are rising in 2026, not falling. H100 rental rates hit a low near $2.79/GPU-hr in mid-2025, then rose about 40% to $2.35/hr on one-year contracts by March 2026 and $2.82/hr on the spot-contract composite by April, driven by a Blackwell supply crunch (CoWoS packaging and HBM capacity) colliding with inference demand that now exceeds training demand. Expect H100/H200 rates to stay elevated or climb further through Q4 2026 as booked capacity clears; expect a real ease only once TSMC's CoWoS gap narrows toward its projected 10% by year-end and Blackwell Ultra volume lands in early 2027. Older cards (A100, V100) keep getting cheaper the whole time — that part of the trend hasn't reversed.

Are cloud GPU prices going up or down right now?

Both, depending which GPU. Frontier hardware (H100, H200, B200) is getting more expensive because supply is constrained and demand from inference workloads is growing faster than new capacity comes online. Older hardware (A100, V100) keeps getting cheaper because it's being displaced down the stack as newer chips arrive — exactly the pattern you'd expect from a market where the top of the line is scarce and everything below it is a hand-me-down.

Here's the H100 rental index SemiAnalysis has tracked since 2023, based on its own historical data:

PeriodH100 rental rate ($/GPU-hr)
2H 2023 (peak)$6.62
1Q 2024$5.78
3Q 2024$4.44
1Q 2025$3.60
2Q–Jul 2025 (trough)$3.34 → $2.79
Jan 2026~$2.80
Mar 2026 (1-yr contract)$2.35
Apr 2026 (spot-contract composite)$2.82

Source: SemiAnalysis GPU rental price index, historical series as published, and SemiAnalysis, "The Great GPU Shortage" for the Oct 2025–Mar 2026 one-year-contract figures. The two series use slightly different baskets (spot-contract composite vs. pure 1-year contract), which is why the exact numbers don't line up month to month — the direction is what matters, and both show the same reversal, in the same window, which is the opposite of what most 2024-era "GPU prices always fall" commentary assumed.

Why: what's actually constraining supply

CoWoS packaging is the bottleneck, not wafer starts. NVIDIA has reportedly secured around 60% of TSMC's CoWoS (chip-on-wafer-on-substrate) advanced packaging capacity — still not enough to clear order volume (wccftech, citing TSMC capacity data). TrendForce reports the CoWoS supply-demand gap narrowing from roughly 20% to a projected 10% by the end of 2026, with TSMC's own monthly CoWoS capacity climbing toward 120,000–140,000 wafers (TrendForce, June 2026). Translation: the gap is closing, but it isn't closed, and "closing" still means every Blackwell GPU shipped this year came out of a capacity-constrained pipe.

HBM memory is the second constraint. SK Hynix holds roughly 62% of the HBM market and is targeting HBM4 mass production in Q3 2026; Samsung is targeting about 50% capacity expansion in 2026 while ramping its own HBM4 later than SK Hynix; Micron is targeting 15,000 wafers/month dedicated to HBM4 by year-end (Tom's Hardware, HBM roadmaps). Every H100, H200, and B200 needs HBM stacked on the package — you can't build the GPU faster than you can build the memory that ships with it.

Demand isn't just training anymore — it's inference, and inference is now the bigger line item. Gartner's most recent AI-optimized IaaS forecast puts 55% of 2026 spending on inference versus 45% on training — the first year inference has overtaken training — with global inference cloud spend at $23.3B against $19B for training, inside a category Gartner expects to grow 96% to $42B this year (Gartner, August 10, 2026). That matters for pricing because inference demand is stickier and more distributed than training demand — it doesn't pause between model runs, and it runs on both frontier chips (for large models) and older cards (for smaller, cheaper-to-serve ones), which pulls on the whole stack at once instead of concentrating pressure on the newest SKU.

The revenue and survey data both back the demand story. NVIDIA's most recently reported quarter (Q1 FY2027, ended April 26, 2026) posted Data Center revenue of $75.2B, up 92% year-over-year, with guidance of $91.0B for the following quarter, a number that explicitly excludes any China Data Center compute revenue (NVIDIA Q1 FY2027 results). NVIDIA's next report lands August 26, 2026 — three days after this update — and we'll fold in fresh numbers next pass rather than guess at them here. More directly, SemiAnalysis's own survey work found on-demand GPU rental capacity "sold out across all GPU types," with roughly half the providers they surveyed completely out of stock on H100/H200 nodes, Blackwell lead times stretching into June–July 2026, and market-wide capacity coming online through August–September 2026 already booked (SemiAnalysis, "The Great GPU Shortage").

What GPUs actually cost right now, by generation

Rates below are on-demand, per-GPU, pulled directly from each provider's own pricing page on August 23, 2026. Where a provider only sells by the node, we show the per-node rate and the GPU count so you can see the basis — don't treat a per-node number as a per-GPU price.

GPULow observed ($/GPU-hr)High observed ($/GPU-hr)Example sources
H100$1.99 (RunPod PCIe, Community)$6.88 (AWS p5.4xlarge, 1x H100)RunPod, AWS via Vantage
H200$3.59 (RunPod, Community)$6.31 (CoreWeave, $50.44/8-GPU node)RunPod, CoreWeave
B200$5.98 (RunPod, Community)$9.86 (Lambda, 16-GPU cluster)RunPod, Lambda
A100 80GB$1.19 (RunPod PCIe, Community)$3.43 (AWS p4de.24xlarge, per-GPU)RunPod, AWS via Vantage
L40S$0.79 (RunPod, Community)$2.25 (CoreWeave, $18.00/8-GPU node)RunPod, CoreWeave

The spread inside a single generation is often wider than the spread between generations — an H100 on RunPod's Community tier ($1.99–$2.89/hr) can be cheaper than an A100 on AWS ($3.43/GPU-hr on the 80GB p4de instance). Generation is a starting point, not the whole answer; the live GPU index and marketplace show what's actually available right now rather than a list price. For the underlying hardware specs behind these numbers, see H100 and A100 — H100 SXM ships with 80GB HBM3 at 3.35 TB/s of bandwidth, per NVIDIA's own datasheet (nvidia.com/h100); A100 ships with 80GB HBM2e at up to 2.04 TB/s (nvidia.com/a100).

The hand-me-down effect: why older GPUs keep getting cheaper

While H100/H200/B200 prices climb, A100 and V100 pricing keeps sliding — the same shortage working in the opposite direction on older silicon. When frontier capacity is scarce, older cards get displaced onto smaller inference jobs and price-sensitive users, and providers compete harder on what's left. Used A100 80GB cards have reportedly fallen to a $3,000–$4,000 secondary-market range, roughly a 60% decline off launch pricing, while used V100 cards sit around $2,000–$3,000, about a 70% decline over five years (Hashrate Index, "The Used GPU Market"; treat these as illustrative industry reporting, not an audited index — we couldn't find a second source corroborating the exact ranges). The pattern isn't new: Azure has reportedly kept V100 fleets in service roughly 7.5 years after launch, and T4 cards were still renting around $0.15/hr seven-plus years after release, per the same reporting. Today's frontier GPU is next cycle's cheap workhorse — the "shortage" headline is a shortage of the newest thing, not of compute generally.

What would prove this call wrong

This is a dated, falsifiable call, not a hedge. We're saying H100/H200/B200 rental rates stay flat-to-up through Q4 2026, with real relief only after TSMC's CoWoS gap closes toward its projected 10% (year-end 2026 target) and Blackwell Ultra (GB300) volume actually ships at scale, which analyst estimates put in early 2027 rather than 2026 (wccftech).

This is wrong if any of the following happens and prices fall anyway: the SemiAnalysis index drops back toward its mid-2025 trough before Q1 2027; a majority of surveyed providers report open (not sold-out) H100/H200 capacity before Q4 2026; or NVIDIA's August 26, 2026 earnings call describes Blackwell supply as no longer constrained. We'll check all three on the next update.

Frequently asked questions

Are cloud GPU prices going up in 2026?

Yes, for frontier GPUs. H100 rental rates bottomed around $2.79/GPU-hr in mid-2025 and rose roughly 40% to $2.35/hr on one-year contracts by March 2026, with the broader spot-contract composite at $2.82/hr by April 2026, per SemiAnalysis's rental index. Older GPUs like A100 and V100 are still getting cheaper over the same period.

Why is there a GPU shortage in 2026?

Two supply-side bottlenecks: TSMC's CoWoS advanced packaging capacity, of which NVIDIA has reportedly secured about 60%, and HBM memory supply from SK Hynix, Samsung, and Micron, none of which reach HBM4 mass production until Q3 2026 at the earliest. Demand-side, inference workloads overtook training as the larger share of AI cloud spending in 2026, adding sustained pressure on top of training demand that never went away.

When will GPU prices drop?

Based on current supply data, meaningful relief is unlikely before TSMC's CoWoS gap narrows to its projected 10% by end of 2026 and Blackwell Ultra ships at volume, which industry estimates place in early 2027. Until then, expect H100/H200/B200 rental rates to hold near current levels or rise further, while A100/V100 rates keep falling.

Will GPU prices drop in 2027?

That's the earliest realistic window based on current packaging and memory capacity projections, but it depends on demand growth not outpacing the new supply. If inference demand keeps growing at anywhere near its 2026 rate, added capacity may get absorbed rather than translate into lower rental rates — this is the single biggest uncertainty in the whole forecast.

What's the cheapest GPU generation to rent right now?

Older datacenter cards (A100, V100) and mid-tier inference cards (L40S) are the cheapest per hour and the ones getting cheaper over time, not more expensive. On-demand L40S rates start around $0.79/GPU-hr on RunPod as of August 2026. If your workload doesn't need H100-class throughput, that's where the actual savings are this year.

Where to check current rates yourself

Every number in the table above ages the moment it's published. For live pricing, check the GPU availability index and marketplace, and the pricing page for cost by workload type; H100 and A100 carry the full spec sheets behind the numbers above.

This arguably strengthens the portability argument: in a market where the cheapest available H100 today might be sold out tomorrow and a better rate opens up on a different provider next week, being able to act on a price change — stop a box, keep the checkpoint and environment, restore somewhere else — matters more when prices are moving than when they're flat.

Sources

#gpu pricing#gpu shortage#h100 pricing#cloud gpu#market trends#blackwell
Ready when you are

Stop paying for
idle GPUs.

Sign up in 60 seconds. Pay only for the GPU minutes you actually use.

© 2026 Aquanode. All rights reserved.

All trademarks, logos and brand names are the property of their respective owners.