NVIDIA L40 GPU: Specs, VRAM, Price & Benchmarks (2026)
The NVIDIA L40 (48GB GDDR6 with ECC) launched in 2022. This guide covers its full specs, VRAM, SXM/PCIe differences where they apply, AI performance, and live per-GPU rental pricing on Aquanode.
Short answer on cost: the L40 rents from $0.759 per GPU per hour on Aquanode. Out of capacity right now.
How much VRAM does the L40 have?
The L40 has 48GB GDDR6 with ECC, with 864 GB/s of peak memory bandwidth.
What fits in 48GB GDDR6 with ECC of VRAM
| Model | Precision | Fits? |
|---|---|---|
| Llama 3 8B | FP16 | roughly 16GB. Fits with room for a real KV cache. |
| Mixtral 8x7B | INT4 | ~24GB. Fits on a single card. |
| Llama 3.1 70B | INT4 | ~35-40GB. Fits, but with little headroom for context. |
| Llama 3.1 70B | FP16 | ~140GB. Does not fit; needs 3+ cards. |
Approximate, based on published parameter counts and standard bytes-per-parameter rules of thumb (FP16 ≈ 2 bytes/param, INT4 ≈ 0.5-0.6 bytes/param). Real footprint also depends on KV-cache size and framework overhead.
What can the L40 run?
Popular open models from small to frontier scale, with the memory each needs and how many L40 cards (48GB GDDR6 with ECC each) that takes.
| Model | As published | FP8 | INT4 |
|---|---|---|---|
| Qwen/Qwen3-8B 8.2B | BF16: ~18.3 GB, 1 GPU | FP8: ~9.2 GB, 1 GPU | INT4: ~4.6 GB, 1 GPU |
| Qwen/Qwen2.5-14B-Instruct 14.8B | BF16: ~33 GB, 1 GPU | FP8: ~16.5 GB, 1 GPU | INT4: ~8.3 GB, 1 GPU |
| Qwen/Qwen3-32B 32.8B | BF16: ~73.2 GB, 2 GPUs | FP8: ~36.6 GB, 1 GPU | INT4: ~18.3 GB, 1 GPU |
| Qwen/Qwen-72B 72.3B | BF16: ~162 GB, 4 GPUs | FP8: ~80.8 GB, 2 GPUs | INT4: ~40.4 GB, 1 GPU |
| MiniMaxAI/MiniMax-M2.7 228.7B | FP8: ~256 GB, 6 GPUs | – | INT4: ~128 GB, 3 GPUs |
| deepseek-ai/DeepSeek-R1 684.5B | FP8: ~765 GB, 16 GPUs | – | INT4: ~383 GB, 8 GPUs |
Estimates: weights at the stated precision plus a flat 20% for KV cache and overhead, at a moderate context length. A dash means the precision is not offered for that model (it is already published at that size). INT4 needs a published quantized checkpoint. Open any model for a per-GPU breakdown, or use the L40 VRAM calculator.
L40 VRAM calculator: check which models fit in its memory at each precision.
All models that fit in 48 GB: the open models whose weights and overhead fit, at native, FP8 and INT4 precision.
L40 specs
| Architecture | NVIDIA, launched 2022 |
| VRAM | 48GB GDDR6 with ECC |
| Memory bandwidth | 864 GB/s |
| FP16 / BF16 tensor throughput | 181.05 TFLOPS (peak, dense) |
| FP8 tensor throughput | 362 TFLOPS (peak, dense) |
| TDP | 300W |
| Form factor | PCIe, dual-slot, passive |
Specs sourced from the vendor's public datasheet/product page. See the source.
Related reading: The NVIDIA Inception program, explained, Google Colab alternatives for dedicated GPU access, Free GPU credits for students and researchers, and RunPod volume disk vs. network volume, compared.
GPU Glossary: What is VRAM?, Tensor Cores, CUDA Cores, TFLOPS
L40 AI performance
The passively-cooled data-center sibling of the L40S: same 48GB of ECC GDDR6 and the same 864 GB/s, with FP8 tensor cores for quantized serving. It's a reasonable card for 7-13B inference and fine-tuning in a rack that can't take SXM parts, and the ECC matters if a long run can't tolerate a silent bit-flip.
- Dense FP16/BF16 tensor throughput: 181.05 TFLOPS
- Dense FP8 tensor throughput: 362 TFLOPS
- Memory bandwidth: 864 GB/s
Roughly half the L40S's dense tensor throughput on the same memory system, so it's the slower card at the same VRAM, and with no NVLink, multi-card training falls back to PCIe. GDDR6 bandwidth also caps large-batch serving well before compute does.
See how it stacks up against other cards in the GPU benchmarks and specs table.
The cheapest L40 offer right now ($0.759/GPU/hr) is about 16% below the market median of $0.902/GPU/hr. Out of capacity right now.
L40 price: what does it cost?
Renting. The cheapest current on-demand rate for the L40 on Aquanode is $0.850/GPU/hr. Live rates range from $0.759 to $1.11 per GPU per hour, with a median of $0.902/GPU/hr. Billed by the offer's own terms; the table below shows every live rate. Out of capacity right now.
| Region | $/GPU/hr | Available | VRAM | vCPU | RAM |
|---|---|---|---|---|---|
| – | $0.759 | 0 | 48 GB | – | – |
| Des Moines, Us | $0.850 | 4 | 48 GB | 12.5 | 72 GB |
| Canada | $1.11 | 1 | 48 GB | 28 | 58 GB |
L40 price history
Aquanode stores one snapshot of its GPU price index per UTC day. For the L40 that is 5 days so far, 2026-10-06 to 2026-10-10, so this is a short history, not a long-run trend. The lowest per-GPU rate went from $0.759 on 2026-10-06 to $0.742 on 2026-10-10.
| Day (UTC) | Lowest per GPU hour | Median per GPU hour | Data-center lowest per GPU hour | Offers |
|---|---|---|---|---|
| 2026-10-10 | $0.742 | $0.902 | $0.850 | 32 |
| 2026-10-09 | $0.759 | $0.902 | $0.850 | 26 |
| 2026-10-08 | $0.759 | $0.902 | $0.850 | 30 |
| 2026-10-07 | $0.759 | $0.902 | $0.850 | 30 |
| 2026-10-06 | $0.759 | $0.902 | $0.850 | 30 |
Each row is the stored daily snapshot of the live index, copied as recorded. A dash means that day stored no figure. The same series is available as JSON and summarised in the monthly GPU price report.
How this price is calculated
All prices on this page are normalized to a per-GPU hourly rate using each offer's authoritative GPU count, so that raw price is divided by the number of GPUs it actually covers; some offers report price as already per-GPU, so those are used as-listed. An offer with a missing, zero, or invalid GPU count is excluded entirely rather than published at a guessed rate.
No offers were excluded from this snapshot for a missing or invalid price. No offers were dropped as price outliers in this snapshot.
Only the cheapest qualifying offer per provider is shown in the table above. This page regenerates at most once per hour.
Compare the L40 with other GPUs
Compare L40 with
- B200 vs L40
- H200 vs L40
- H100 vs L40
- A100 vs L40
- DGX A100 vs L40
- L40 vs V100
- AMD MI300X vs L40
- L40 vs L40S
- A40 vs L40
- L40 vs RTX A6000
- L40 vs RTX A5000
- L40 vs RTX A4000
- L4 vs L40
- L40 vs T4
- L40 vs RTX 6000 Ada
- L40 vs RTX PRO 6000
- L40 vs RTX PRO 6000 WS
- L40 vs RTX PRO 6000 SE
- L40 vs RTX PRO 5000
- L40 vs RTX 5090
- L40 vs RTX 5080
- L40 vs RTX 5070 Ti
- L40 vs RTX 5070
- L40 vs RTX 5060 Ti
- L40 vs RTX 4090
- L40 vs RTX 4080 Super
- L40 vs RTX 4080
- L40 vs RTX 4070 Ti
- L40 vs RTX 4070 Super
- L40 vs RTX 4070
- L40 vs RTX 4060 Ti
- L40 vs RTX 3090
- L40 vs RTX 3080
- L40 vs RTX 3070
- L40 vs RTX 3060
- B300 vs L40
- AMD MI355X vs L40
- AMD MI325X vs L40
- H100 NVL vs L40
- L40 vs RTX PRO 4500
- L40 vs RTX PRO 4500 SE
- L40 vs RTX PRO 4000
- L40 vs RTX 5880 Ada
- L40 vs RTX 5000 Ada
- L40 vs RTX 4000 Ada
- L40 vs RTX 4000 SFF Ada
- L40 vs RTX 2000 Ada
- L40 vs RTX 4070 Ti Super
- L40 vs RTX A4500
- L40 vs RTX 2080 Ti
- L40 vs RTX 3080 Ti
- L40 vs RTX 3070 Ti
- L40 vs RTX 3060 Ti
- GTX 1080 Ti vs L40
- L40 vs RTX 4060
- L40 vs RTX 5060
- GTX 1660 Super vs L40
- L40 vs RTX 2060
- L40 vs Radeon RX 9070 XT
- L40 vs Radeon RX 6700 XT
- L40 vs RTX 6000
- A16 vs L40
- L40 vs P40
- L40 vs P4
- L40 vs Quadro P2000
- L40 vs Quadro M4000
Get notified when the price drops
GPU supply moves hourly. Tell us what you're waiting for and we'll email you when a matching offer appears across any provider we track.
Rent a L40 on Aquanode
- On-demand instances from $0.759/GPU/hr, billed by the provider's own terms, with no hardware procurement or long-term commitment.
- 28 live offers across 3 regions today.
- Set a price/availability alert above to hear the moment a cheaper or newly-available L40 offer appears.
- Compare every L40 offer side by side, or browse the full multi-provider GPU marketplace.
Good for
The passively-cooled data-center sibling of the L40S: same 48GB of ECC GDDR6 and the same 864 GB/s, with FP8 tensor cores for quantized serving. It's a reasonable card for 7-13B inference and fine-tuning in a rack that can't take SXM parts, and the ECC matters if a long run can't tolerate a silent bit-flip.
Not good for
Roughly half the L40S's dense tensor throughput on the same memory system, so it's the slower card at the same VRAM, and with no NVLink, multi-card training falls back to PCIe. GDDR6 bandwidth also caps large-batch serving well before compute does.
L40 FAQs
How much VRAM does the L40 have?
The L40 has 48GB GDDR6 with ECC, with 864 GB/s of peak memory bandwidth.
What is the NVIDIA L40?
The NVIDIA L40 is a GPU released in 2022, with 48GB GDDR6 with ECC of memory and a 300W power envelope. See the full spec table above for interconnect, form factor and tensor-throughput details.
How much does it cost to rent a L40?
Live L40 rental prices currently range from $0.759 to $1.11 per GPU per hour, with a median of $0.902 per GPU per hour. Out of capacity right now.
What's the cheapest L40 rate?
The lowest current L40 rate on Aquanode is $0.759 per GPU per hour in . Out of capacity right now.
What is the L40 good for?
The passively-cooled data-center sibling of the L40S: same 48GB of ECC GDDR6 and the same 864 GB/s, with FP8 tensor cores for quantized serving. It's a reasonable card for 7-13B inference and fine-tuning in a rack that can't take SXM parts, and the ECC matters if a long run can't tolerate a silent bit-flip.
What are the L40's limitations?
Roughly half the L40S's dense tensor throughput on the same memory system, so it's the slower card at the same VRAM, and with no NVLink, multi-card training falls back to PCIe. GDDR6 bandwidth also caps large-batch serving well before compute does.
Is renting cheaper than buying?
Renting avoids the upfront hardware cost and lets you match spend to actual usage. A rented L40 at $0.759/hr only costs money while it's running, whereas buying ties up capital in hardware that keeps depreciating whether it's in use or not. Which is cheaper depends on how continuously you'd run it; short or bursty workloads usually favor renting. Out of capacity right now.
How is the L40 price calculated?
All prices are normalized to a per-GPU hourly rate using each offer's authoritative GPU count, which the raw price is divided by; some offers report price as already per-GPU. Offers whose price can't be safely normalized, or whose rate is an extreme outlier against the rest of the market, are excluded.
L40 price by region
Related guides
Other models in the same generation, then the rest of the GPU index.