NVIDIA L40 GPU: Specs, VRAM, Price & Benchmarks (2026)

The NVIDIA L40 (48GB GDDR6 with ECC) launched in 2022. This guide covers its full specs, VRAM, SXM/PCIe differences where they apply, AI performance, and live per-GPU rental pricing on Aquanode.

Short answer on cost: the L40 rents from $0.759 per GPU per hour on Aquanode. Out of capacity right now.

How much VRAM does the L40 have?

The L40 has 48GB GDDR6 with ECC, with 864 GB/s of peak memory bandwidth.

What fits in 48GB GDDR6 with ECC of VRAM

ModelPrecisionFits?
Llama 3 8BFP16roughly 16GB. Fits with room for a real KV cache.
Mixtral 8x7BINT4~24GB. Fits on a single card.
Llama 3.1 70BINT4~35-40GB. Fits, but with little headroom for context.
Llama 3.1 70BFP16~140GB. Does not fit; needs 3+ cards.

Approximate, based on published parameter counts and standard bytes-per-parameter rules of thumb (FP16 ≈ 2 bytes/param, INT4 ≈ 0.5-0.6 bytes/param). Real footprint also depends on KV-cache size and framework overhead.

What can the L40 run?

Popular open models from small to frontier scale, with the memory each needs and how many L40 cards (48GB GDDR6 with ECC each) that takes.

ModelAs publishedFP8INT4
Qwen/Qwen3-8B 8.2BBF16: ~18.3 GB, 1 GPUFP8: ~9.2 GB, 1 GPUINT4: ~4.6 GB, 1 GPU
Qwen/Qwen2.5-14B-Instruct 14.8BBF16: ~33 GB, 1 GPUFP8: ~16.5 GB, 1 GPUINT4: ~8.3 GB, 1 GPU
Qwen/Qwen3-32B 32.8BBF16: ~73.2 GB, 2 GPUsFP8: ~36.6 GB, 1 GPUINT4: ~18.3 GB, 1 GPU
Qwen/Qwen-72B 72.3BBF16: ~162 GB, 4 GPUsFP8: ~80.8 GB, 2 GPUsINT4: ~40.4 GB, 1 GPU
MiniMaxAI/MiniMax-M2.7 228.7BFP8: ~256 GB, 6 GPUs–INT4: ~128 GB, 3 GPUs
deepseek-ai/DeepSeek-R1 684.5BFP8: ~765 GB, 16 GPUs–INT4: ~383 GB, 8 GPUs

Estimates: weights at the stated precision plus a flat 20% for KV cache and overhead, at a moderate context length. A dash means the precision is not offered for that model (it is already published at that size). INT4 needs a published quantized checkpoint. Open any model for a per-GPU breakdown, or use the L40 VRAM calculator.

L40 VRAM calculator: check which models fit in its memory at each precision.

All models that fit in 48 GB: the open models whose weights and overhead fit, at native, FP8 and INT4 precision.

$0.759/GPU/hrOut of capacity right now
Lowest / GPU / hr
$0.902/GPU/hr
Median / GPU / hr
$1.11/GPU/hr
p90 / GPU / hr
28
Live offers
Last updated: 2026-10-10 21:03:10 UTCRefreshes hourly0 offer(s) excluded from this snapshot

L40 specs

ArchitectureNVIDIA, launched 2022
VRAM48GB GDDR6 with ECC
Memory bandwidth864 GB/s
FP16 / BF16 tensor throughput181.05 TFLOPS (peak, dense)
FP8 tensor throughput362 TFLOPS (peak, dense)
TDP300W
Form factorPCIe, dual-slot, passive

Specs sourced from the vendor's public datasheet/product page. See the source.

Related reading: The NVIDIA Inception program, explained, Google Colab alternatives for dedicated GPU access, Free GPU credits for students and researchers, and RunPod volume disk vs. network volume, compared.

GPU Glossary: What is VRAM?, Tensor Cores, CUDA Cores, TFLOPS

L40 AI performance

The passively-cooled data-center sibling of the L40S: same 48GB of ECC GDDR6 and the same 864 GB/s, with FP8 tensor cores for quantized serving. It's a reasonable card for 7-13B inference and fine-tuning in a rack that can't take SXM parts, and the ECC matters if a long run can't tolerate a silent bit-flip.

  • Dense FP16/BF16 tensor throughput: 181.05 TFLOPS
  • Dense FP8 tensor throughput: 362 TFLOPS
  • Memory bandwidth: 864 GB/s

Roughly half the L40S's dense tensor throughput on the same memory system, so it's the slower card at the same VRAM, and with no NVLink, multi-card training falls back to PCIe. GDDR6 bandwidth also caps large-batch serving well before compute does.

See how it stacks up against other cards in the GPU benchmarks and specs table.

The cheapest L40 offer right now ($0.759/GPU/hr) is about 16% below the market median of $0.902/GPU/hr. Out of capacity right now.

L40 price: what does it cost?

Renting. The cheapest current on-demand rate for the L40 on Aquanode is $0.850/GPU/hr. Live rates range from $0.759 to $1.11 per GPU per hour, with a median of $0.902/GPU/hr. Billed by the offer's own terms; the table below shows every live rate. Out of capacity right now.

Region$/GPU/hrAvailableVRAMvCPURAM
–$0.759048 GB––
Des Moines, Us$0.850448 GB12.572 GB
Canada$1.11148 GB2858 GB

L40 price history

Aquanode stores one snapshot of its GPU price index per UTC day. For the L40 that is 5 days so far, 2026-10-06 to 2026-10-10, so this is a short history, not a long-run trend. The lowest per-GPU rate went from $0.759 on 2026-10-06 to $0.742 on 2026-10-10.

Day (UTC)Lowest per GPU hourMedian per GPU hourData-center lowest per GPU hourOffers
2026-10-10$0.742$0.902$0.85032
2026-10-09$0.759$0.902$0.85026
2026-10-08$0.759$0.902$0.85030
2026-10-07$0.759$0.902$0.85030
2026-10-06$0.759$0.902$0.85030

Each row is the stored daily snapshot of the live index, copied as recorded. A dash means that day stored no figure. The same series is available as JSON and summarised in the monthly GPU price report.

How this price is calculated

All prices on this page are normalized to a per-GPU hourly rate using each offer's authoritative GPU count, so that raw price is divided by the number of GPUs it actually covers; some offers report price as already per-GPU, so those are used as-listed. An offer with a missing, zero, or invalid GPU count is excluded entirely rather than published at a guessed rate.

No offers were excluded from this snapshot for a missing or invalid price. No offers were dropped as price outliers in this snapshot.

Only the cheapest qualifying offer per provider is shown in the table above. This page regenerates at most once per hour.

Compare the L40 with other GPUs

Compare L40 with

Get notified when the price drops

GPU supply moves hourly. Tell us what you're waiting for and we'll email you when a matching offer appears across any provider we track.

One email per matching alert. Unsubscribe any time.

Rent a L40 on Aquanode

  • On-demand instances from $0.759/GPU/hr, billed by the provider's own terms, with no hardware procurement or long-term commitment.
  • 28 live offers across 3 regions today.
  • Set a price/availability alert above to hear the moment a cheaper or newly-available L40 offer appears.
  • Compare every L40 offer side by side, or browse the full multi-provider GPU marketplace.

Good for

The passively-cooled data-center sibling of the L40S: same 48GB of ECC GDDR6 and the same 864 GB/s, with FP8 tensor cores for quantized serving. It's a reasonable card for 7-13B inference and fine-tuning in a rack that can't take SXM parts, and the ECC matters if a long run can't tolerate a silent bit-flip.

Not good for

Roughly half the L40S's dense tensor throughput on the same memory system, so it's the slower card at the same VRAM, and with no NVLink, multi-card training falls back to PCIe. GDDR6 bandwidth also caps large-batch serving well before compute does.

L40 FAQs

How much VRAM does the L40 have?

The L40 has 48GB GDDR6 with ECC, with 864 GB/s of peak memory bandwidth.

What is the NVIDIA L40?

The NVIDIA L40 is a GPU released in 2022, with 48GB GDDR6 with ECC of memory and a 300W power envelope. See the full spec table above for interconnect, form factor and tensor-throughput details.

How much does it cost to rent a L40?

Live L40 rental prices currently range from $0.759 to $1.11 per GPU per hour, with a median of $0.902 per GPU per hour. Out of capacity right now.

What's the cheapest L40 rate?

The lowest current L40 rate on Aquanode is $0.759 per GPU per hour in . Out of capacity right now.

What is the L40 good for?

The passively-cooled data-center sibling of the L40S: same 48GB of ECC GDDR6 and the same 864 GB/s, with FP8 tensor cores for quantized serving. It's a reasonable card for 7-13B inference and fine-tuning in a rack that can't take SXM parts, and the ECC matters if a long run can't tolerate a silent bit-flip.

What are the L40's limitations?

Roughly half the L40S's dense tensor throughput on the same memory system, so it's the slower card at the same VRAM, and with no NVLink, multi-card training falls back to PCIe. GDDR6 bandwidth also caps large-batch serving well before compute does.

Is renting cheaper than buying?

Renting avoids the upfront hardware cost and lets you match spend to actual usage. A rented L40 at $0.759/hr only costs money while it's running, whereas buying ties up capital in hardware that keeps depreciating whether it's in use or not. Which is cheaper depends on how continuously you'd run it; short or bursty workloads usually favor renting. Out of capacity right now.

How is the L40 price calculated?

All prices are normalized to a per-GPU hourly rate using each offer's authoritative GPU count, which the raw price is divided by; some offers report price as already per-GPU. Offers whose price can't be safely normalized, or whose rate is an extreme outlier against the rest of the market, are excluded.

L40 price by region

Related guides

Other models in the same generation, then the rest of the GPU index.

Submit the job. Everything after that is ours.

Sign up in 60 seconds. Pay for the GPU minutes you actually use.

© 2026 Aquanode. All rights reserved.

All trademarks, logos and brand names are the property of their respective owners.