The RTX 5090 is the most-listed GPU on our marketplace right now, and it has the widest price range of anything we carry. Same card, same hour: $0.158 at the floor, $2.004 at the top. Run it eight hours a day for a month and that's the difference between a $38 bill and a $481 one. This is what actually drives that gap, what the card buys you over a 4090, and the one setup detail that wastes an afternoon if you don't know it before you rent.
TL;DR: 80 live offers, median $0.704/GPU/hr, floor $0.158, ceiling $2.004. Multi-GPU nodes are consistently cheaper per GPU than single cards, and the expensive listings cluster in high-electricity-cost regions. Versus a 4090 you get 32GB instead of 24GB and — the part people miss — native FP4 tensor cores, which Ada does not have. You also need PyTorch 2.7+ on CUDA 12.8; older stacks have no wheels that run on this card.
What the market actually looks like
Every number here is from one snapshot of our own marketplace feed on August 4, 2026 — 80 live RTX 5090 offers across three providers (Vast.ai 77, Akash 2, SimplePod 1), normalized to dollars per GPU per hour.
| $/GPU/hr | Monthly at 8h/day | |
|---|---|---|
| Cheapest | $0.158 | $38 |
| 25th percentile | $0.501 | $120 |
| Median | $0.704 | $169 |
| 75th percentile | $0.869 | $209 |
| 90th percentile | $1.038 | $249 |
| Most expensive | $2.004 | $481 |
For comparison, RunPod's own RTX 5090 page lists $0.69/hr on Community Cloud and $0.99/hr on Secure Cloud (checked August 4, 2026). That's roughly our median and our 75th percentile respectively — which is the honest read: a large single-vendor platform prices near the middle of the market, and the tails are where marketplace shopping pays.
Two patterns in the spread are worth knowing before you shop.
Multi-GPU nodes are cheaper per GPU. The four cheapest Vast.ai listings are 4-GPU and 2-GPU boxes; the most expensive are single cards. If your workload can use more than one GPU, or you're willing to rent a 2-card box and use one, the per-GPU rate drops noticeably.
Geography is doing real work. Supply is genuinely global — South Korea (7 offers), Spain (6), then California, Hong Kong, Belgium, Alberta, Estonia and Norway at 4 each. The $2.004 outlier is in Norway; the $1.476 is in the UK. Consumer GPU hosting is an electricity-price business, and it shows.
The live version of this table is on the RTX 5090 price page, which is where to look before acting on numbers that were true on one August afternoon.
What you're actually renting
NVIDIA's own RTX 5090 spec page against the 4090:
| RTX 5090 | RTX 4090 | |
|---|---|---|
| Architecture | Blackwell | Ada Lovelace |
| VRAM | 32 GB GDDR7 | 24 GB GDDR6X |
| Memory bus | 512-bit | 384-bit |
| Memory bandwidth | 1,792 GB/s | ~1,008 GB/s |
| CUDA cores | 21,760 | 16,384 |
| Tensor cores | 5th gen | 4th gen |
| Total graphics power | 575 W | 450 W |
| Launched | Jan 30, 2025 | Oct 12, 2022 |
| MSRP at launch | $1,999 | $1,599 |
One honesty note on that table: I could not pull the 4090's bandwidth figure as text from NVIDIA's own page (it renders client-side), so the ~1,008 GB/s is the widely-cited number and is arithmetically consistent with a 384-bit bus at 21 Gbps. Every other cell came off NVIDIA's pages directly.
The 8GB question, answered with arithmetic instead of marketing
The upgrade everyone quotes is 24GB to 32GB. Here's what it actually changes, using roughly 2 bytes per parameter at FP16, 1 at INT8, and about 0.6 at INT4, with usable VRAM after overhead around 20–22GB on a 4090 and 28–30GB on a 5090.
- 7–8B models fit at FP16 on both. No change.
- 13B models need ~26GB at FP16 — doesn't fit a 4090, fits a 5090 with modest context. At INT8 both are fine.
- 30–34B models at INT8 need ~32GB, which only the 5090 reaches, and barely. At INT4 (~18–20GB) both are comfortable.
- Mixtral 8x7B at INT4 is ~26GB: comfortable on a 5090, tight to impossible on a 4090 once KV cache is added.
- 70B models at INT4 are ~40–42GB. Neither card gets there. The extra 8GB does not put a 70B model on a single consumer GPU, and anyone implying otherwise is selling you something.
So the real dividing line isn't 70B. It's that 30–47B-class models go from "not really" to "comfortably," and image and video models get room to run at higher precision instead of being forced down.
The part people miss: FP4 is a Blackwell-only capability
The memory difference gets all the attention, and it's the less interesting half. From NVIDIA's own Blackwell architecture whitepaper:
"The RTX Blackwell Tensor Cores support FP16, BF16, TF32, INT8, and Hopper's FP8 Transformer Engine. RTX Blackwell adds new support for FP4 and FP6 Tensor Core operations, and the new Second-Generation FP8 Transformer Engine."
The 4090's 4th-gen tensor cores have no native FP4. That's an architectural gap, not a spec-sheet gradient. NVIDIA's own worked example in the same document:
"With a GeForce RTX 4090 with FP16, the FLUX.dev model can generate images in 15 seconds with 30 steps. With a GeForce RTX 5090 with FP4, images can be generated in just over five seconds."
And on memory, from the same whitepaper: FLUX.dev at FP16 "requires over 23GB of VRAM" — right at the 4090's ceiling with no headroom for batching — while at FP4 it "requires less than 10GB."
If you're doing diffusion work and your toolchain supports FP4, that combination is a much stronger argument for the 5090 than the 8GB is. If your toolchain doesn't support FP4 yet, you're paying for a capability you can't reach, and a cheap 4090 may be the better rental.
The setup detail that wastes an afternoon
Blackwell consumer cards are sm_120, and older builds have no kernels for it.
PyTorch 2.7.0 is the first stable release with Blackwell support, listed in its release notes as a beta feature alongside CUDA 12.8 support. You need the CUDA 12.8 wheels:
pip install torch --index-url https://download.pytorch.org/whl/cu128
Anything older than 2.7 has no prebuilt wheel that runs on this card. If you rent a 5090 and land on an image pinned to an older PyTorch, it will fail in a way that looks like a broken box rather than a version problem — which is exactly the sort of thing that eats an hour of billed GPU time before you work out what's wrong.
The ecosystem around PyTorch lagged behind it. Reports of needing to build flash-attention from source, or compile bitsandbytes for Blackwell, were common after 2.7 landed. I could not find official release-notes pages pinning down the exact first supported versions for those two the way PyTorch's notes do, so treat it as: check your specific dependencies against sm_120 before you rent for a long block, and test on an hourly instance first.
When a 5090 is the right rental
Rent one when you're doing diffusion or video generation and can use FP4, when you're serving a 13–47B model that doesn't fit a 4090 comfortably, or when you want the fastest single consumer card available and the price you found is near the market floor rather than the ceiling.
Rent a 4090 instead when your models fit in 24GB, your stack doesn't use FP4, and you can find one cheap. It's an older card at a lower price and for a lot of workloads the 5090 buys you nothing you can use.
Rent a datacenter card instead when you need more than 32GB on one device, when you need ECC or multi-GPU interconnect that consumer cards don't provide, or when you have procurement requirements that consumer hardware doesn't satisfy. Our cross-provider price analysis has the datacenter comparison — the short version is that A100 and H100 prices vary far less than consumer cards do, so there's less to gain from shopping and more to gain from availability.
Why the spread is only useful if you can move
A 12x range on the same card is a large number, and it's worth exactly as much as your ability to act on it. If taking the $0.158 box instead of the $0.704 one means an afternoon reinstalling CUDA 12.8, rebuilding your environment and re-pulling model weights, then the saving is notional. That's especially true here, because a Blackwell setup takes more fiddling than most — and having done it once, you really don't want to do it again on the next box.
That's the problem we work on: capturing the whole machine so that moving to a cheaper or more available one isn't a rebuild. The mechanics are in moving a GPU workload to another cloud provider and snapshot and restore across providers.
FAQ
How much does it cost to rent an RTX 5090? Across 80 live offers on August 4, 2026, the median was $0.704 per GPU per hour, with a floor of $0.158 and a ceiling of $2.004. RunPod's own page lists $0.69/hr Community and $0.99/hr Secure. Live figures are on the RTX 5090 price page.
Is an RTX 5090 worth it over a 4090 for AI? If you use FP4 for diffusion work, or your model needs 24–32GB, yes. If your models fit in 24GB and your stack doesn't use FP4, a cheaper 4090 often makes more sense.
Can an RTX 5090 run a 70B model? Not on its own at usable quality. A 70B model at INT4 needs roughly 40–42GB, above the 5090's 32GB. You'd need multiple GPUs or CPU offloading.
Why is the same 5090 so much cheaper on some listings? Multi-GPU nodes have lower per-GPU rates than single cards, and hosting costs vary hugely by region — the most expensive listings in this snapshot were single cards in Norway and the UK.
What CUDA version does an RTX 5090 need?
PyTorch 2.7.0 with CUDA 12.8 wheels is the first stable combination with Blackwell (sm_120) support. Older PyTorch builds have no wheels that run on the card.
Does an RTX 5090 support FP8 and FP4? Yes. Blackwell's 5th-gen tensor cores add native FP4 and FP6 plus a second-generation FP8 Transformer Engine, per NVIDIA's architecture whitepaper. The 4090 does not have native FP4.
The short version
The 5090 is a genuinely better AI card than the 4090, but for a more specific reason than the one on the box: FP4 support changes what diffusion work costs, while the extra 8GB moves a narrow band of model sizes from awkward to fine. Neither gets you to 70B.
And the price you pay has more to do with which listing you click than which card you picked. A 12x range on identical silicon is not a market failure, it's an unpriced difference in electricity, node size and host. Shop the tails.
About the author
I'm Ansh Saxena. I work on the infrastructure layer under rented GPU boxes, mostly on making a machine's whole state portable so that acting on a price difference doesn't cost you a rebuild. The pricing figures here come from a public feed you can query yourself rather than a number you have to take on trust.
Sources
- Aquanode marketplace feed, snapshot 2026-08-04: 80 RTX 5090 offers across Vast.ai (77), Akash (2) and SimplePod (1). Live at /gpu/rtx-5090 and /marketplace.
- NVIDIA GeForce RTX 5090 — 32GB GDDR7, 512-bit bus, 1,792 GB/s, 21,760 CUDA cores, 5th-gen tensor cores, 575W.
- NVIDIA GeForce RTX 4090 — 24GB GDDR6X, 384-bit bus, 16,384 CUDA cores, 450W.
- NVIDIA RTX Blackwell GPU Architecture whitepaper — FP4/FP6 tensor core support, second-gen FP8 Transformer Engine, FLUX.dev VRAM and generation-time figures.
- PyTorch 2.7.0 release notes — Blackwell architecture support (beta), CUDA 12.8.
- RunPod RTX 5090 pricing, checked 2026-08-04 — $0.69/hr Community Cloud, $0.99/hr Secure Cloud.