NVIDIA L40S GPU: Specs, VRAM, Price & Benchmarks (2026)

The NVIDIA L40S (48GB GDDR6 with ECC) launched in 2023. This guide covers its full specs, VRAM, SXM/PCIe differences where they apply, AI performance, and live per-GPU rental pricing on Aquanode.

Short answer on cost: the L40S rents from $0.698 per GPU per hour on Aquanode.

How much VRAM does the L40S have?

The L40S has 48GB GDDR6 with ECC, with 864 GB/s of peak memory bandwidth.

  • VRAM: L40S 48GB GDDR6 with ECC vs RTX A6000 48GB GDDR6 with ECC

What fits in 48GB GDDR6 with ECC of VRAM

ModelPrecisionFits?
Llama 3 8BFP16~16GB. Fits with room for a large batch and KV cache.
Llama 3.1 70BFP16~140GB. Does not fit; needs multiple cards.
Llama 3.1 70BINT4~35-40GB. Fits on a single card.
Mistral 7BFP8~8GB. Fits comfortably, leaves room for high-concurrency serving.

Approximate, based on published parameter counts and standard bytes-per-parameter rules of thumb (FP16 ≈ 2 bytes/param, INT4 ≈ 0.5-0.6 bytes/param). Real footprint also depends on KV-cache size and framework overhead.

What can the L40S run?

Popular open models from small to frontier scale, with the memory each needs and how many L40S cards (48GB GDDR6 with ECC each) that takes.

ModelAs publishedFP8INT4
Qwen/Qwen3-8B 8.2BBF16: ~18.3 GB, 1 GPUFP8: ~9.2 GB, 1 GPUINT4: ~4.6 GB, 1 GPU
Qwen/Qwen2.5-14B-Instruct 14.8BBF16: ~33 GB, 1 GPUFP8: ~16.5 GB, 1 GPUINT4: ~8.3 GB, 1 GPU
Qwen/Qwen3-32B 32.8BBF16: ~73.2 GB, 2 GPUsFP8: ~36.6 GB, 1 GPUINT4: ~18.3 GB, 1 GPU
Qwen/Qwen-72B 72.3BBF16: ~162 GB, 4 GPUsFP8: ~80.8 GB, 2 GPUsINT4: ~40.4 GB, 1 GPU
MiniMaxAI/MiniMax-M2.7 228.7BFP8: ~256 GB, 6 GPUs–INT4: ~128 GB, 3 GPUs
deepseek-ai/DeepSeek-R1 684.5BFP8: ~765 GB, 16 GPUs–INT4: ~383 GB, 8 GPUs

Estimates: weights at the stated precision plus a flat 20% for KV cache and overhead, at a moderate context length. A dash means the precision is not offered for that model (it is already published at that size). INT4 needs a published quantized checkpoint. Open any model for a per-GPU breakdown, or use the L40S VRAM calculator.

L40S VRAM calculator: check which models fit in its memory at each precision.

All models that fit in 48 GB: the open models whose weights and overhead fit, at native, FP8 and INT4 precision.

$0.698/GPU/hr
Lowest / GPU / hr
$1.20/GPU/hr
Median / GPU / hr
$2.57/GPU/hr
p90 / GPU / hr
52
Live offers
Last updated: 2026-10-10 21:02:29 UTCRefreshes hourly4 offer(s) excluded from this snapshot

L40S specs

ArchitectureNVIDIA, launched 2023
VRAM48GB GDDR6 with ECC
Memory bandwidth864 GB/s
FP16 / BF16 tensor throughput362 TFLOPS (peak, dense)
FP8 tensor throughput733 TFLOPS (peak, dense)
TDP350W
Form factorPCIe, dual-slot, air-cooled

Specs sourced from the vendor's public datasheet/product page. See the source.

Related reading: The best GPUs for AI, ranked.

GPU Glossary: What is VRAM?, Tensor Cores, CUDA Cores, TFLOPS

L40S AI performance

A PCIe, air-cooled inference and fine-tuning card that doesn't need SXM/NVLink server infrastructure. 48GB is enough for most 7-13B models in FP16 and larger models in quantized form, and its FP8 tensor cores make it a solid throughput-per-dollar choice for serving.

  • Dense FP16/BF16 tensor throughput: 362 TFLOPS
  • Dense FP8 tensor throughput: 733 TFLOPS
  • Memory bandwidth: 864 GB/s

GDDR6 bandwidth (864 GB/s) is well below HBM parts like H100 (3.35 TB/s), so it's memory-bandwidth-bound on large-batch or long-context serving well before it's compute-bound, and no NVLink means no fast GPU-to-GPU path for multi-card training.

See how it stacks up against other cards in the GPU benchmarks and specs table.

The cheapest L40S offer right now ($0.698/GPU/hr) is about 42% below the market median of $1.20/GPU/hr.

L40S price: what does it cost?

Buying. $7,500-$10,000 (commonly quoted street price (no fixed retail; OEM channel)).

Renting. The cheapest current on-demand rate for the L40S on Aquanode is $1.07/GPU/hr. Live rates range from $0.698 to $2.57 per GPU per hour, with a median of $1.20/GPU/hr. Billed by the offer's own terms; the table below shows every live rate.

Region$/GPU/hrAvailableVRAMvCPURAM
Türkiye, Tr$0.698245 GB12126 GB
–$0.869048 GB24251 GB
Kansas City, Us$1.07148 GB1272 GB
Finland$1.701648 GB832 GB
Finland$1.78148 GB2060 GB
Virginia, Us$2.05045 GB432 GB

L40S price history

Aquanode stores one snapshot of its GPU price index per UTC day. For the L40S that is 5 days so far, 2026-10-06 to 2026-10-10, so this is a short history, not a long-run trend. The lowest per-GPU rate went from $0.869 on 2026-10-06 to $0.726 on 2026-10-10.

Day (UTC)Lowest per GPU hourMedian per GPU hourData-center lowest per GPU hourOffers
2026-10-10$0.726$1.20$1.0750
2026-10-09$0.869$1.20$1.0743
2026-10-08$0.869$1.20$1.0740
2026-10-07$0.869$1.20$1.0750
2026-10-06$0.869$1.20$1.0739

Each row is the stored daily snapshot of the live index, copied as recorded. A dash means that day stored no figure. The same series is available as JSON and summarised in the monthly GPU price report.

How this price is calculated

All prices on this page are normalized to a per-GPU hourly rate using each offer's authoritative GPU count, so that raw price is divided by the number of GPUs it actually covers; some offers report price as already per-GPU, so those are used as-listed. An offer with a missing, zero, or invalid GPU count is excluded entirely rather than published at a guessed rate.

No offers were excluded from this snapshot for a missing or invalid price. 4 offer(s) were dropped as high-side price outliers (more than 3x the model's median rate).

Only the cheapest qualifying offer per provider is shown in the table above. This page regenerates at most once per hour.

L40S vs RTX A6000: how do they compare?

  • VRAM: L40S 48GB GDDR6 with ECC vs RTX A6000 48GB GDDR6 with ECC
  • Memory bandwidth: 864 GB/s vs 768 GB/s

On Aquanode right now, RTX A6000 starts at $0.363/GPU/hr against L40S's $0.698/GPU/hr, about 48% less.

Full L40S vs RTX A6000 price comparison

Compare the L40S with other GPUs

Compare L40S with

Get notified when the price drops

GPU supply moves hourly. Tell us what you're waiting for and we'll email you when a matching offer appears across any provider we track.

One email per matching alert. Unsubscribe any time.

Rent a L40S on Aquanode

  • On-demand instances from $0.698/GPU/hr, billed by the provider's own terms, with no hardware procurement or long-term commitment.
  • 52 live offers across 5 regions today.
  • Set a price/availability alert above to hear the moment a cheaper or newly-available L40S offer appears.
  • Compare every L40S offer side by side, or browse the full multi-provider GPU marketplace.

Good for

A PCIe, air-cooled inference and fine-tuning card that doesn't need SXM/NVLink server infrastructure. 48GB is enough for most 7-13B models in FP16 and larger models in quantized form, and its FP8 tensor cores make it a solid throughput-per-dollar choice for serving.

Not good for

GDDR6 bandwidth (864 GB/s) is well below HBM parts like H100 (3.35 TB/s), so it's memory-bandwidth-bound on large-batch or long-context serving well before it's compute-bound, and no NVLink means no fast GPU-to-GPU path for multi-card training.

L40S FAQs

How much VRAM does the L40S have?

The L40S has 48GB GDDR6 with ECC, with 864 GB/s of peak memory bandwidth.

What is the NVIDIA L40S?

The NVIDIA L40S is a GPU released in 2023, with 48GB GDDR6 with ECC of memory and a 350W power envelope. See the full spec table above for interconnect, form factor and tensor-throughput details.

How much does it cost to rent a L40S?

Live L40S rental prices currently range from $0.698 to $2.57 per GPU per hour, with a median of $1.20 per GPU per hour.

What's the cheapest L40S rate?

The lowest current L40S rate on Aquanode is $0.698 per GPU per hour in Türkiye, Tr.

How does the L40S compare to the RTX A6000?

VRAM: L40S 48GB GDDR6 with ECC vs RTX A6000 48GB GDDR6 with ECC Memory bandwidth: 864 GB/s vs 768 GB/s On Aquanode right now, RTX A6000 starts at $0.363/GPU/hr against L40S's $0.698/GPU/hr, about 48% less. See the full L40S vs RTX A6000 comparison for a shared-provider price breakdown.

What is the L40S good for?

A PCIe, air-cooled inference and fine-tuning card that doesn't need SXM/NVLink server infrastructure. 48GB is enough for most 7-13B models in FP16 and larger models in quantized form, and its FP8 tensor cores make it a solid throughput-per-dollar choice for serving.

What are the L40S's limitations?

GDDR6 bandwidth (864 GB/s) is well below HBM parts like H100 (3.35 TB/s), so it's memory-bandwidth-bound on large-batch or long-context serving well before it's compute-bound, and no NVLink means no fast GPU-to-GPU path for multi-card training.

Is renting cheaper than buying?

Renting avoids the upfront hardware cost and lets you match spend to actual usage. A rented L40S at $0.698/hr only costs money while it's running, whereas buying ties up capital in hardware that keeps depreciating whether it's in use or not. Buying outright runs $7,500-$10,000 (commonly quoted street price (no fixed retail; OEM channel)). Which is cheaper depends on how continuously you'd run it; short or bursty workloads usually favor renting.

How is the L40S price calculated?

All prices are normalized to a per-GPU hourly rate using each offer's authoritative GPU count, which the raw price is divided by; some offers report price as already per-GPU. Offers whose price can't be safely normalized, or whose rate is an extreme outlier against the rest of the market, are excluded.

L40S price by region

Related guides

Other models in the same generation, then the rest of the GPU index.

Submit the job. Everything after that is ours.

Sign up in 60 seconds. Pay for the GPU minutes you actually use.

© 2026 Aquanode. All rights reserved.

All trademarks, logos and brand names are the property of their respective owners.