NVIDIA L4 GPU: Specs, VRAM, Price & Benchmarks (2026)

The NVIDIA L4 (24GB GDDR6) launched in 2023. This guide covers its full specs, VRAM, SXM/PCIe differences where they apply, AI performance, and live per-GPU rental pricing on Aquanode.

Short answer on cost: the L4 rents from $0.649 per GPU per hour on Aquanode. Out of capacity right now.

How much VRAM does the L4 have?

The L4 has 24GB GDDR6, with 300 GB/s of peak memory bandwidth.

What fits in 24GB GDDR6 of VRAM

ModelPrecisionFits?
Llama 3 8BFP16~16GB. Fits, though with a tight KV-cache budget.
Llama 3 8BINT4~5-6GB. Fits comfortably, leaves room for a real batch size.
Mistral 7BINT4~4-5GB. Fits easily.
Llama 3.1 70BINT4~35-40GB. Does not fit on a single 24GB card.

Approximate, based on published parameter counts and standard bytes-per-parameter rules of thumb (FP16 ≈ 2 bytes/param, INT4 ≈ 0.5-0.6 bytes/param). Real footprint also depends on KV-cache size and framework overhead.

What can the L4 run?

Popular open models from small to frontier scale, with the memory each needs and how many L4 cards (24GB GDDR6 each) that takes.

ModelAs publishedFP8INT4
Qwen/Qwen3-8B 8.2BBF16: ~18.3 GB, 1 GPUFP8: ~9.2 GB, 1 GPUINT4: ~4.6 GB, 1 GPU
Qwen/Qwen2.5-14B-Instruct 14.8BBF16: ~33 GB, 2 GPUsFP8: ~16.5 GB, 1 GPUINT4: ~8.3 GB, 1 GPU
Qwen/Qwen3-32B 32.8BBF16: ~73.2 GB, 4 GPUsFP8: ~36.6 GB, 2 GPUsINT4: ~18.3 GB, 1 GPU
Qwen/Qwen-72B 72.3BBF16: ~162 GB, 7 GPUsFP8: ~80.8 GB, 4 GPUsINT4: ~40.4 GB, 2 GPUs
MiniMaxAI/MiniMax-M2.7 228.7BFP8: ~256 GB, 11 GPUs–INT4: ~128 GB, 6 GPUs
deepseek-ai/DeepSeek-R1 684.5BFP8: ~765 GB, 32 GPUs–INT4: ~383 GB, 16 GPUs

Estimates: weights at the stated precision plus a flat 20% for KV cache and overhead, at a moderate context length. A dash means the precision is not offered for that model (it is already published at that size). INT4 needs a published quantized checkpoint. Open any model for a per-GPU breakdown, or use the L4 VRAM calculator.

L4 VRAM calculator: check which models fit in its memory at each precision.

All models that fit in 24 GB: the open models whose weights and overhead fit, at native, FP8 and INT4 precision.

$0.649/GPU/hrOut of capacity right now
Lowest / GPU / hr
$0.649/GPU/hr
Median / GPU / hr
$0.649/GPU/hr
p90 / GPU / hr
15
Live offers
Last updated: 2026-10-10 21:04:53 UTCRefreshes hourly0 offer(s) excluded from this snapshot

L4 specs

ArchitectureNVIDIA, launched 2023
VRAM24GB GDDR6
Memory bandwidth300 GB/s
TDP72W
Form factorPCIe, 1-slot, low-profile

Specs sourced from the vendor's public datasheet/product page. See the source.

Related reading: The best GPUs for AI, ranked.

GPU Glossary: What is VRAM?, Tensor Cores, CUDA Cores, TFLOPS

L4 AI performance

The lowest-power card on this marketplace (72W, single low-profile slot). Built for dense inference deployments and video/AI pipelines where power and rack density matter more than raw throughput, not for training.

  • Memory bandwidth: 300 GB/s

24GB VRAM and 300 GB/s bandwidth rule out anything beyond small-model inference. A 70B-class model doesn't fit even quantized, and it has no NVLink for multi-card scale-up.

L4 price: what does it cost?

Renting. The cheapest current on-demand rate for the L4 on Aquanode is $0.649/GPU/hr. Live rates range from $0.649 to $0.649 per GPU per hour, with a median of $0.649/GPU/hr. Billed by the offer's own terms; the table below shows every live rate. Out of capacity right now.

Region$/GPU/hrAvailableVRAMvCPURAM
Romania$0.649024 GB655 GB

L4 price history

Aquanode stores one snapshot of its GPU price index per UTC day. For the L4 that is 5 days so far, 2026-10-06 to 2026-10-10, so this is a short history, not a long-run trend. The lowest per-GPU rate went from $0.478 on 2026-10-06 to $0.500 on 2026-10-10.

Day (UTC)Lowest per GPU hourMedian per GPU hourData-center lowest per GPU hourOffers
2026-10-10$0.500$0.649$0.64916
2026-10-09$0.478$0.649$0.64914
2026-10-08$0.405$0.539$0.53915
2026-10-07$0.478$0.539$0.53920
2026-10-06$0.478$0.539$0.53913

Each row is the stored daily snapshot of the live index, copied as recorded. A dash means that day stored no figure. The same series is available as JSON and summarised in the monthly GPU price report.

How this price is calculated

All prices on this page are normalized to a per-GPU hourly rate using each offer's authoritative GPU count, so that raw price is divided by the number of GPUs it actually covers; some offers report price as already per-GPU, so those are used as-listed. An offer with a missing, zero, or invalid GPU count is excluded entirely rather than published at a guessed rate.

No offers were excluded from this snapshot for a missing or invalid price. No offers were dropped as price outliers in this snapshot.

Only the cheapest qualifying offer per provider is shown in the table above. This page regenerates at most once per hour.

Compare the L4 with other GPUs

Compare L4 with

Get notified when the price drops

GPU supply moves hourly. Tell us what you're waiting for and we'll email you when a matching offer appears across any provider we track.

One email per matching alert. Unsubscribe any time.

Rent a L4 on Aquanode

  • On-demand instances from $0.649/GPU/hr, billed by the provider's own terms, with no hardware procurement or long-term commitment.
  • 15 live offers across 1 region today.
  • Set a price/availability alert above to hear the moment a cheaper or newly-available L4 offer appears.
  • Compare every L4 offer side by side, or browse the full multi-provider GPU marketplace.

Good for

The lowest-power card on this marketplace (72W, single low-profile slot). Built for dense inference deployments and video/AI pipelines where power and rack density matter more than raw throughput, not for training.

Not good for

24GB VRAM and 300 GB/s bandwidth rule out anything beyond small-model inference. A 70B-class model doesn't fit even quantized, and it has no NVLink for multi-card scale-up.

L4 FAQs

How much VRAM does the L4 have?

The L4 has 24GB GDDR6, with 300 GB/s of peak memory bandwidth.

What is the NVIDIA L4?

The NVIDIA L4 is a GPU released in 2023, with 24GB GDDR6 of memory and a 72W power envelope. See the full spec table above for interconnect, form factor and tensor-throughput details.

How much does it cost to rent a L4?

Live L4 rental prices currently range from $0.649 to $0.649 per GPU per hour, with a median of $0.649 per GPU per hour. Out of capacity right now.

What's the cheapest L4 rate?

The lowest current L4 rate on Aquanode is $0.649 per GPU per hour in Romania. Out of capacity right now.

What is the L4 good for?

The lowest-power card on this marketplace (72W, single low-profile slot). Built for dense inference deployments and video/AI pipelines where power and rack density matter more than raw throughput, not for training.

What are the L4's limitations?

24GB VRAM and 300 GB/s bandwidth rule out anything beyond small-model inference. A 70B-class model doesn't fit even quantized, and it has no NVLink for multi-card scale-up.

Is renting cheaper than buying?

Renting avoids the upfront hardware cost and lets you match spend to actual usage. A rented L4 at $0.649/hr only costs money while it's running, whereas buying ties up capital in hardware that keeps depreciating whether it's in use or not. Which is cheaper depends on how continuously you'd run it; short or bursty workloads usually favor renting. Out of capacity right now.

How is the L4 price calculated?

All prices are normalized to a per-GPU hourly rate using each offer's authoritative GPU count, which the raw price is divided by; some offers report price as already per-GPU. Offers whose price can't be safely normalized, or whose rate is an extreme outlier against the rest of the market, are excluded.

L4 price by region

Related guides

Other models in the same generation, then the rest of the GPU index.

Submit the job. Everything after that is ours.

Sign up in 60 seconds. Pay for the GPU minutes you actually use.

© 2026 Aquanode. All rights reserved.

All trademarks, logos and brand names are the property of their respective owners.