NVIDIA L40S GPU: Specs, VRAM, Price & Benchmarks (2026)
The NVIDIA L40S (48GB GDDR6 with ECC) launched in 2023. This guide covers its full specs, VRAM, SXM/PCIe differences where they apply, AI performance, and live per-GPU rental pricing on Aquanode.
Short answer on cost: the L40S rents from $0.698 per GPU per hour on Aquanode.
How much VRAM does the L40S have?
The L40S has 48GB GDDR6 with ECC, with 864 GB/s of peak memory bandwidth.
- VRAM: L40S 48GB GDDR6 with ECC vs RTX A6000 48GB GDDR6 with ECC
What fits in 48GB GDDR6 with ECC of VRAM
| Model | Precision | Fits? |
|---|---|---|
| Llama 3 8B | FP16 | ~16GB. Fits with room for a large batch and KV cache. |
| Llama 3.1 70B | FP16 | ~140GB. Does not fit; needs multiple cards. |
| Llama 3.1 70B | INT4 | ~35-40GB. Fits on a single card. |
| Mistral 7B | FP8 | ~8GB. Fits comfortably, leaves room for high-concurrency serving. |
Approximate, based on published parameter counts and standard bytes-per-parameter rules of thumb (FP16 ≈ 2 bytes/param, INT4 ≈ 0.5-0.6 bytes/param). Real footprint also depends on KV-cache size and framework overhead.
What can the L40S run?
Popular open models from small to frontier scale, with the memory each needs and how many L40S cards (48GB GDDR6 with ECC each) that takes.
| Model | As published | FP8 | INT4 |
|---|---|---|---|
| Qwen/Qwen3-8B 8.2B | BF16: ~18.3 GB, 1 GPU | FP8: ~9.2 GB, 1 GPU | INT4: ~4.6 GB, 1 GPU |
| Qwen/Qwen2.5-14B-Instruct 14.8B | BF16: ~33 GB, 1 GPU | FP8: ~16.5 GB, 1 GPU | INT4: ~8.3 GB, 1 GPU |
| Qwen/Qwen3-32B 32.8B | BF16: ~73.2 GB, 2 GPUs | FP8: ~36.6 GB, 1 GPU | INT4: ~18.3 GB, 1 GPU |
| Qwen/Qwen-72B 72.3B | BF16: ~162 GB, 4 GPUs | FP8: ~80.8 GB, 2 GPUs | INT4: ~40.4 GB, 1 GPU |
| MiniMaxAI/MiniMax-M2.7 228.7B | FP8: ~256 GB, 6 GPUs | – | INT4: ~128 GB, 3 GPUs |
| deepseek-ai/DeepSeek-R1 684.5B | FP8: ~765 GB, 16 GPUs | – | INT4: ~383 GB, 8 GPUs |
Estimates: weights at the stated precision plus a flat 20% for KV cache and overhead, at a moderate context length. A dash means the precision is not offered for that model (it is already published at that size). INT4 needs a published quantized checkpoint. Open any model for a per-GPU breakdown, or use the L40S VRAM calculator.
L40S VRAM calculator: check which models fit in its memory at each precision.
All models that fit in 48 GB: the open models whose weights and overhead fit, at native, FP8 and INT4 precision.
L40S specs
| Architecture | NVIDIA, launched 2023 |
| VRAM | 48GB GDDR6 with ECC |
| Memory bandwidth | 864 GB/s |
| FP16 / BF16 tensor throughput | 362 TFLOPS (peak, dense) |
| FP8 tensor throughput | 733 TFLOPS (peak, dense) |
| TDP | 350W |
| Form factor | PCIe, dual-slot, air-cooled |
Specs sourced from the vendor's public datasheet/product page. See the source.
Related reading: The best GPUs for AI, ranked.
GPU Glossary: What is VRAM?, Tensor Cores, CUDA Cores, TFLOPS
L40S AI performance
A PCIe, air-cooled inference and fine-tuning card that doesn't need SXM/NVLink server infrastructure. 48GB is enough for most 7-13B models in FP16 and larger models in quantized form, and its FP8 tensor cores make it a solid throughput-per-dollar choice for serving.
- Dense FP16/BF16 tensor throughput: 362 TFLOPS
- Dense FP8 tensor throughput: 733 TFLOPS
- Memory bandwidth: 864 GB/s
GDDR6 bandwidth (864 GB/s) is well below HBM parts like H100 (3.35 TB/s), so it's memory-bandwidth-bound on large-batch or long-context serving well before it's compute-bound, and no NVLink means no fast GPU-to-GPU path for multi-card training.
See how it stacks up against other cards in the GPU benchmarks and specs table.
The cheapest L40S offer right now ($0.698/GPU/hr) is about 42% below the market median of $1.20/GPU/hr.
L40S price: what does it cost?
Buying. $7,500-$10,000 (commonly quoted street price (no fixed retail; OEM channel)).
Renting. The cheapest current on-demand rate for the L40S on Aquanode is $1.07/GPU/hr. Live rates range from $0.698 to $2.57 per GPU per hour, with a median of $1.20/GPU/hr. Billed by the offer's own terms; the table below shows every live rate.
| Region | $/GPU/hr | Available | VRAM | vCPU | RAM |
|---|---|---|---|---|---|
| Türkiye, Tr | $0.698 | 2 | 45 GB | 12 | 126 GB |
| – | $0.869 | 0 | 48 GB | 24 | 251 GB |
| Kansas City, Us | $1.07 | 1 | 48 GB | 12 | 72 GB |
| Finland | $1.70 | 16 | 48 GB | 8 | 32 GB |
| Finland | $1.78 | 1 | 48 GB | 20 | 60 GB |
| Virginia, Us | $2.05 | 0 | 45 GB | 4 | 32 GB |
L40S price history
Aquanode stores one snapshot of its GPU price index per UTC day. For the L40S that is 5 days so far, 2026-10-06 to 2026-10-10, so this is a short history, not a long-run trend. The lowest per-GPU rate went from $0.869 on 2026-10-06 to $0.726 on 2026-10-10.
| Day (UTC) | Lowest per GPU hour | Median per GPU hour | Data-center lowest per GPU hour | Offers |
|---|---|---|---|---|
| 2026-10-10 | $0.726 | $1.20 | $1.07 | 50 |
| 2026-10-09 | $0.869 | $1.20 | $1.07 | 43 |
| 2026-10-08 | $0.869 | $1.20 | $1.07 | 40 |
| 2026-10-07 | $0.869 | $1.20 | $1.07 | 50 |
| 2026-10-06 | $0.869 | $1.20 | $1.07 | 39 |
Each row is the stored daily snapshot of the live index, copied as recorded. A dash means that day stored no figure. The same series is available as JSON and summarised in the monthly GPU price report.
How this price is calculated
All prices on this page are normalized to a per-GPU hourly rate using each offer's authoritative GPU count, so that raw price is divided by the number of GPUs it actually covers; some offers report price as already per-GPU, so those are used as-listed. An offer with a missing, zero, or invalid GPU count is excluded entirely rather than published at a guessed rate.
No offers were excluded from this snapshot for a missing or invalid price. 4 offer(s) were dropped as high-side price outliers (more than 3x the model's median rate).
Only the cheapest qualifying offer per provider is shown in the table above. This page regenerates at most once per hour.
L40S vs RTX A6000: how do they compare?
- VRAM: L40S 48GB GDDR6 with ECC vs RTX A6000 48GB GDDR6 with ECC
- Memory bandwidth: 864 GB/s vs 768 GB/s
On Aquanode right now, RTX A6000 starts at $0.363/GPU/hr against L40S's $0.698/GPU/hr, about 48% less.
Compare the L40S with other GPUs
Compare L40S with
- B200 vs L40S
- H200 vs L40S
- H100 vs L40S
- A100 vs L40S
- DGX A100 vs L40S
- L40S vs V100
- AMD MI300X vs L40S
- L40 vs L40S
- A40 vs L40S
- L40S vs RTX A6000
- L40S vs RTX A5000
- L40S vs RTX A4000
- L4 vs L40S
- L40S vs T4
- L40S vs RTX 6000 Ada
- L40S vs RTX PRO 6000
- L40S vs RTX PRO 6000 WS
- L40S vs RTX PRO 6000 SE
- L40S vs RTX PRO 5000
- L40S vs RTX 5090
- L40S vs RTX 5080
- L40S vs RTX 5070 Ti
- L40S vs RTX 5070
- L40S vs RTX 5060 Ti
- L40S vs RTX 4090
- L40S vs RTX 4080 Super
- L40S vs RTX 4080
- L40S vs RTX 4070 Ti
- L40S vs RTX 4070 Super
- L40S vs RTX 4070
- L40S vs RTX 4060 Ti
- L40S vs RTX 3090
- L40S vs RTX 3080
- L40S vs RTX 3070
- L40S vs RTX 3060
- B300 vs L40S
- AMD MI355X vs L40S
- AMD MI325X vs L40S
- H100 NVL vs L40S
- L40S vs RTX PRO 4500
- L40S vs RTX PRO 4500 SE
- L40S vs RTX PRO 4000
- L40S vs RTX 5880 Ada
- L40S vs RTX 5000 Ada
- L40S vs RTX 4000 Ada
- L40S vs RTX 4000 SFF Ada
- L40S vs RTX 2000 Ada
- L40S vs RTX 4070 Ti Super
- L40S vs RTX A4500
- L40S vs RTX 2080 Ti
- L40S vs RTX 3080 Ti
- L40S vs RTX 3070 Ti
- L40S vs RTX 3060 Ti
- GTX 1080 Ti vs L40S
- L40S vs RTX 4060
- L40S vs RTX 5060
- GTX 1660 Super vs L40S
- L40S vs RTX 2060
- L40S vs Radeon RX 9070 XT
- L40S vs Radeon RX 6700 XT
- L40S vs RTX 6000
- A16 vs L40S
- L40S vs P40
- L40S vs P4
- L40S vs Quadro P2000
- L40S vs Quadro M4000
Get notified when the price drops
GPU supply moves hourly. Tell us what you're waiting for and we'll email you when a matching offer appears across any provider we track.
Rent a L40S on Aquanode
- On-demand instances from $0.698/GPU/hr, billed by the provider's own terms, with no hardware procurement or long-term commitment.
- 52 live offers across 5 regions today.
- Set a price/availability alert above to hear the moment a cheaper or newly-available L40S offer appears.
- Compare every L40S offer side by side, or browse the full multi-provider GPU marketplace.
Good for
A PCIe, air-cooled inference and fine-tuning card that doesn't need SXM/NVLink server infrastructure. 48GB is enough for most 7-13B models in FP16 and larger models in quantized form, and its FP8 tensor cores make it a solid throughput-per-dollar choice for serving.
Not good for
GDDR6 bandwidth (864 GB/s) is well below HBM parts like H100 (3.35 TB/s), so it's memory-bandwidth-bound on large-batch or long-context serving well before it's compute-bound, and no NVLink means no fast GPU-to-GPU path for multi-card training.
L40S FAQs
How much VRAM does the L40S have?
The L40S has 48GB GDDR6 with ECC, with 864 GB/s of peak memory bandwidth.
What is the NVIDIA L40S?
The NVIDIA L40S is a GPU released in 2023, with 48GB GDDR6 with ECC of memory and a 350W power envelope. See the full spec table above for interconnect, form factor and tensor-throughput details.
How much does it cost to rent a L40S?
Live L40S rental prices currently range from $0.698 to $2.57 per GPU per hour, with a median of $1.20 per GPU per hour.
What's the cheapest L40S rate?
The lowest current L40S rate on Aquanode is $0.698 per GPU per hour in Türkiye, Tr.
How does the L40S compare to the RTX A6000?
VRAM: L40S 48GB GDDR6 with ECC vs RTX A6000 48GB GDDR6 with ECC Memory bandwidth: 864 GB/s vs 768 GB/s On Aquanode right now, RTX A6000 starts at $0.363/GPU/hr against L40S's $0.698/GPU/hr, about 48% less. See the full L40S vs RTX A6000 comparison for a shared-provider price breakdown.
What is the L40S good for?
A PCIe, air-cooled inference and fine-tuning card that doesn't need SXM/NVLink server infrastructure. 48GB is enough for most 7-13B models in FP16 and larger models in quantized form, and its FP8 tensor cores make it a solid throughput-per-dollar choice for serving.
What are the L40S's limitations?
GDDR6 bandwidth (864 GB/s) is well below HBM parts like H100 (3.35 TB/s), so it's memory-bandwidth-bound on large-batch or long-context serving well before it's compute-bound, and no NVLink means no fast GPU-to-GPU path for multi-card training.
Is renting cheaper than buying?
Renting avoids the upfront hardware cost and lets you match spend to actual usage. A rented L40S at $0.698/hr only costs money while it's running, whereas buying ties up capital in hardware that keeps depreciating whether it's in use or not. Buying outright runs $7,500-$10,000 (commonly quoted street price (no fixed retail; OEM channel)). Which is cheaper depends on how continuously you'd run it; short or bursty workloads usually favor renting.
How is the L40S price calculated?
All prices are normalized to a per-GPU hourly rate using each offer's authoritative GPU count, which the raw price is divided by; some offers report price as already per-GPU. Offers whose price can't be safely normalized, or whose rate is an extreme outlier against the rest of the market, are excluded.
L40S price by region
Related guides
Other models in the same generation, then the rest of the GPU index.