L40S vs H100: cloud rental cost compared

Short answer: pick the H100 when you're training, fine-tuning, or serving at a scale that's bandwidth-bound. Pick the L40S when you're running inference at moderate scale and want a lower hourly rate with a card still built for AI workloads.

You can rent both by the hour, no purchase required. H100 from $1.99/GPU/hr across 8 providers; L40S from $0.790/GPU/hr across 5 providers. The L40S is 60% cheaper at the entry rate.

H100
Good for training and serving frontier-scale LLMs
$1.99/hr
Lowest / GPU
$3.35/hr
Median / GPU
80 GB
VRAM
Hopper
Architecture
L40S
Good for fine-tuning mid-size models and high-throughput inference
$0.790/hr
Lowest / GPU
$1.09/hr
Median / GPU
48 GB
VRAM
Ada Lovelace
Architecture
Last updated: 2026-09-19 06:19:44 UTCRefreshes hourly

Specs side by side

Spec
H100
L40S
VRAM
80 GB
48 GB
Architecture
Hopper
Ada Lovelace
Compute capability
9.0
8.9
Precisions in hardware
FP32, FP16, BF16, FP8, INT4
FP32, FP16, BF16, FP8, INT4
Interconnect
SXM5
PCIe
Lowest $/GPU/hr
$1.99
$0.790
Median $/GPU/hr
$3.35
$1.09
Providers
8
5
Regions
8
5

Price by provider

Which should you pick: L40S or H100?

The H100 has roughly 4x the memory bandwidth of the L40S, 3.35 TB/s of HBM3 against 864 GB/s of GDDR6 (NVIDIA product pages, Aug 2026), and is the SXM data-center accelerator, while the L40S is the more affordable data-center-AI card built for inference density rather than peak training throughput. Pick the H100 when you're training, fine-tuning, or serving at a scale that's bandwidth-bound; pick the L40S when you're running inference at moderate scale and want a lower hourly rate with a card still built for AI workloads. On price the L40S undercuts the H100 by 60% at the entry rate ($0.790 vs $1.99/GPU/hr). Worth weighing if the deciding factor above isn't a hard requirement for your job.

Decision criteria

Memory bandwidth
H100 SXM: 3.35 TB/s HBM3. L40S: 864 GB/s GDDR6 (NVIDIA product pages, Aug 2026).
Positioning
H100 is the flagship SXM training/inference accelerator; L40S is a lower-cost, inference-optimized data-center card.
VRAM
80 GB (H100) vs 48 GB (L40S), live from current listings.
Availability
8 provider(s) list the H100 vs 5 for the L40S.
Entry price
$1.99/GPU/hr (H100) vs $0.790/GPU/hr (L40S).

Pick the H100 when

  • You're training or fine-tuning at scale
  • You're bandwidth-bound on large-batch serving
  • You need the flagship data-center SXM accelerator

Pick the L40S when

  • You're running inference at moderate scale
  • You want a lower hourly rate on a card still built for AI
  • L40S has better availability where you're deploying

Price-performance

At list rates, 1,000 GPU-hours costs $790 on the L40S against $1,990 on the H100. Neither NVIDIA nor Aquanode publishes a workload-normalized $/token or $/epoch figure for this pair, so $/GPU-hour below is the only apples-to-apples number. A faster card can still cost less per finished job even at a higher hourly rate.

H100 vs L40S: common questions

Is the H100 or the L40S cheaper to rent?

On Aquanode's live marketplace the H100 starts at $1.99/GPU/hr (median $3.35/GPU/hr across 8 providers) and the L40S starts at $0.790/GPU/hr (median $1.09/GPU/hr across 5 providers). The L40S is the cheaper of the two at the entry rate, by 60%.

What is the difference between the H100 and the L40S?

The H100 has 80 GB of VRAM against the L40S's 48 GB; the H100 is a Hopper part and the L40S is Ada Lovelace (compute capability 9.0 vs 8.9); both run BF16, FP8, INT4 workloads in hardware. On price, the H100 lists from $1.99/GPU/hr and the L40S from $0.790/GPU/hr.

Which cloud providers offer the H100 and the L40S?

8 providers list the H100 (RunPod, Massed Compute, HyperStack, Jarvislabs, Akash, Vast.ai, Verda and Nebius) and 5 list the L40S (RunPod, Vast.ai, Massed Compute, Verda and Nebius), across 8 and 5 regions respectively. Both are available from RunPod, Massed Compute, Vast.ai, Verda and Nebius.

How much does 1,000 GPU-hours cost on the H100 vs the L40S?

At the lowest rates listed today, 1,000 GPU-hours costs $1,990 on the H100 and $790 on the L40S, a difference of $1,200 for the same runtime. Rates are per GPU per hour and update hourly.

Is the L40S a good budget alternative to the H100 for inference?

For moderate-scale inference, often yes. It has FP8 support like the H100, just at 864 GB/s of bandwidth against the H100's 3.35 TB/s (NVIDIA, Aug 2026). Today the L40S lists from $0.790/GPU/hr against the H100's $1.99/GPU/hr on Aquanode, so it's worth testing your actual throughput before committing at scale.

How this comparison is calculated

Every price is normalized to a per-GPU hourly rate using the same pipeline as every other pricing surface on Aquanode, and only the cheapest qualifying offer per provider is shown. A dash means that provider does not currently list that GPU.

Architecture, compute capability and supported precisions come from each vendor's published datasheet for that generation, not from the marketplace feed. This page regenerates at most once per hour.

Ready when you are

Submit the job.
A dead GPU doesn't end it.

Sign up in 60 seconds. Pay for the GPU minutes you actually use.

© 2026 Aquanode. All rights reserved.

All trademarks, logos and brand names are the property of their respective owners.