NVIDIA L4 GPU: Specs, VRAM, Price & Benchmarks (2026)
The NVIDIA L4 (24GB GDDR6) launched in 2023. This guide covers its full specs, VRAM, SXM/PCIe differences where they apply, AI performance, and live per-GPU rental pricing on Aquanode.
Short answer on cost: the L4 rents from $0.649 per GPU per hour on Aquanode. Out of capacity right now.
How much VRAM does the L4 have?
The L4 has 24GB GDDR6, with 300 GB/s of peak memory bandwidth.
What fits in 24GB GDDR6 of VRAM
| Model | Precision | Fits? |
|---|---|---|
| Llama 3 8B | FP16 | ~16GB. Fits, though with a tight KV-cache budget. |
| Llama 3 8B | INT4 | ~5-6GB. Fits comfortably, leaves room for a real batch size. |
| Mistral 7B | INT4 | ~4-5GB. Fits easily. |
| Llama 3.1 70B | INT4 | ~35-40GB. Does not fit on a single 24GB card. |
Approximate, based on published parameter counts and standard bytes-per-parameter rules of thumb (FP16 ≈ 2 bytes/param, INT4 ≈ 0.5-0.6 bytes/param). Real footprint also depends on KV-cache size and framework overhead.
What can the L4 run?
Popular open models from small to frontier scale, with the memory each needs and how many L4 cards (24GB GDDR6 each) that takes.
| Model | As published | FP8 | INT4 |
|---|---|---|---|
| Qwen/Qwen3-8B 8.2B | BF16: ~18.3 GB, 1 GPU | FP8: ~9.2 GB, 1 GPU | INT4: ~4.6 GB, 1 GPU |
| Qwen/Qwen2.5-14B-Instruct 14.8B | BF16: ~33 GB, 2 GPUs | FP8: ~16.5 GB, 1 GPU | INT4: ~8.3 GB, 1 GPU |
| Qwen/Qwen3-32B 32.8B | BF16: ~73.2 GB, 4 GPUs | FP8: ~36.6 GB, 2 GPUs | INT4: ~18.3 GB, 1 GPU |
| Qwen/Qwen-72B 72.3B | BF16: ~162 GB, 7 GPUs | FP8: ~80.8 GB, 4 GPUs | INT4: ~40.4 GB, 2 GPUs |
| MiniMaxAI/MiniMax-M2.7 228.7B | FP8: ~256 GB, 11 GPUs | – | INT4: ~128 GB, 6 GPUs |
| deepseek-ai/DeepSeek-R1 684.5B | FP8: ~765 GB, 32 GPUs | – | INT4: ~383 GB, 16 GPUs |
Estimates: weights at the stated precision plus a flat 20% for KV cache and overhead, at a moderate context length. A dash means the precision is not offered for that model (it is already published at that size). INT4 needs a published quantized checkpoint. Open any model for a per-GPU breakdown, or use the L4 VRAM calculator.
L4 VRAM calculator: check which models fit in its memory at each precision.
All models that fit in 24 GB: the open models whose weights and overhead fit, at native, FP8 and INT4 precision.
L4 specs
| Architecture | NVIDIA, launched 2023 |
| VRAM | 24GB GDDR6 |
| Memory bandwidth | 300 GB/s |
| TDP | 72W |
| Form factor | PCIe, 1-slot, low-profile |
Specs sourced from the vendor's public datasheet/product page. See the source.
Related reading: The best GPUs for AI, ranked.
GPU Glossary: What is VRAM?, Tensor Cores, CUDA Cores, TFLOPS
L4 AI performance
The lowest-power card on this marketplace (72W, single low-profile slot). Built for dense inference deployments and video/AI pipelines where power and rack density matter more than raw throughput, not for training.
- Memory bandwidth: 300 GB/s
24GB VRAM and 300 GB/s bandwidth rule out anything beyond small-model inference. A 70B-class model doesn't fit even quantized, and it has no NVLink for multi-card scale-up.
L4 price: what does it cost?
Renting. The cheapest current on-demand rate for the L4 on Aquanode is $0.649/GPU/hr. Live rates range from $0.649 to $0.649 per GPU per hour, with a median of $0.649/GPU/hr. Billed by the offer's own terms; the table below shows every live rate. Out of capacity right now.
| Region | $/GPU/hr | Available | VRAM | vCPU | RAM |
|---|---|---|---|---|---|
| Romania | $0.649 | 0 | 24 GB | 6 | 55 GB |
L4 price history
Aquanode stores one snapshot of its GPU price index per UTC day. For the L4 that is 5 days so far, 2026-10-06 to 2026-10-10, so this is a short history, not a long-run trend. The lowest per-GPU rate went from $0.478 on 2026-10-06 to $0.500 on 2026-10-10.
| Day (UTC) | Lowest per GPU hour | Median per GPU hour | Data-center lowest per GPU hour | Offers |
|---|---|---|---|---|
| 2026-10-10 | $0.500 | $0.649 | $0.649 | 16 |
| 2026-10-09 | $0.478 | $0.649 | $0.649 | 14 |
| 2026-10-08 | $0.405 | $0.539 | $0.539 | 15 |
| 2026-10-07 | $0.478 | $0.539 | $0.539 | 20 |
| 2026-10-06 | $0.478 | $0.539 | $0.539 | 13 |
Each row is the stored daily snapshot of the live index, copied as recorded. A dash means that day stored no figure. The same series is available as JSON and summarised in the monthly GPU price report.
How this price is calculated
All prices on this page are normalized to a per-GPU hourly rate using each offer's authoritative GPU count, so that raw price is divided by the number of GPUs it actually covers; some offers report price as already per-GPU, so those are used as-listed. An offer with a missing, zero, or invalid GPU count is excluded entirely rather than published at a guessed rate.
No offers were excluded from this snapshot for a missing or invalid price. No offers were dropped as price outliers in this snapshot.
Only the cheapest qualifying offer per provider is shown in the table above. This page regenerates at most once per hour.
Compare the L4 with other GPUs
Compare L4 with
- B200 vs L4
- H200 vs L4
- H100 vs L4
- A100 vs L4
- DGX A100 vs L4
- L4 vs V100
- AMD MI300X vs L4
- L4 vs L40S
- L4 vs L40
- A40 vs L4
- L4 vs RTX A6000
- L4 vs RTX A5000
- L4 vs RTX A4000
- L4 vs T4
- L4 vs RTX 6000 Ada
- L4 vs RTX PRO 6000
- L4 vs RTX PRO 6000 WS
- L4 vs RTX PRO 6000 SE
- L4 vs RTX PRO 5000
- L4 vs RTX 5090
- L4 vs RTX 5080
- L4 vs RTX 5070 Ti
- L4 vs RTX 5070
- L4 vs RTX 5060 Ti
- L4 vs RTX 4090
- L4 vs RTX 4080 Super
- L4 vs RTX 4080
- L4 vs RTX 4070 Ti
- L4 vs RTX 4070 Super
- L4 vs RTX 4070
- L4 vs RTX 4060 Ti
- L4 vs RTX 3090
- L4 vs RTX 3080
- L4 vs RTX 3070
- L4 vs RTX 3060
- B300 vs L4
- AMD MI355X vs L4
- AMD MI325X vs L4
- H100 NVL vs L4
- L4 vs RTX PRO 4500
- L4 vs RTX PRO 4500 SE
- L4 vs RTX PRO 4000
- L4 vs RTX 5880 Ada
- L4 vs RTX 5000 Ada
- L4 vs RTX 4000 Ada
- L4 vs RTX 4000 SFF Ada
- L4 vs RTX 2000 Ada
- L4 vs RTX 4070 Ti Super
- L4 vs RTX A4500
- L4 vs RTX 2080 Ti
- L4 vs RTX 3080 Ti
- L4 vs RTX 3070 Ti
- L4 vs RTX 3060 Ti
- GTX 1080 Ti vs L4
- L4 vs RTX 4060
- L4 vs RTX 5060
- GTX 1660 Super vs L4
- L4 vs RTX 2060
- L4 vs Radeon RX 9070 XT
- L4 vs Radeon RX 6700 XT
- L4 vs RTX 6000
- A16 vs L4
- L4 vs P40
- L4 vs P4
- L4 vs Quadro P2000
- L4 vs Quadro M4000
Get notified when the price drops
GPU supply moves hourly. Tell us what you're waiting for and we'll email you when a matching offer appears across any provider we track.
Rent a L4 on Aquanode
- On-demand instances from $0.649/GPU/hr, billed by the provider's own terms, with no hardware procurement or long-term commitment.
- 15 live offers across 1 region today.
- Set a price/availability alert above to hear the moment a cheaper or newly-available L4 offer appears.
- Compare every L4 offer side by side, or browse the full multi-provider GPU marketplace.
Good for
The lowest-power card on this marketplace (72W, single low-profile slot). Built for dense inference deployments and video/AI pipelines where power and rack density matter more than raw throughput, not for training.
Not good for
24GB VRAM and 300 GB/s bandwidth rule out anything beyond small-model inference. A 70B-class model doesn't fit even quantized, and it has no NVLink for multi-card scale-up.
L4 FAQs
How much VRAM does the L4 have?
The L4 has 24GB GDDR6, with 300 GB/s of peak memory bandwidth.
What is the NVIDIA L4?
The NVIDIA L4 is a GPU released in 2023, with 24GB GDDR6 of memory and a 72W power envelope. See the full spec table above for interconnect, form factor and tensor-throughput details.
How much does it cost to rent a L4?
Live L4 rental prices currently range from $0.649 to $0.649 per GPU per hour, with a median of $0.649 per GPU per hour. Out of capacity right now.
What's the cheapest L4 rate?
The lowest current L4 rate on Aquanode is $0.649 per GPU per hour in Romania. Out of capacity right now.
What is the L4 good for?
The lowest-power card on this marketplace (72W, single low-profile slot). Built for dense inference deployments and video/AI pipelines where power and rack density matter more than raw throughput, not for training.
What are the L4's limitations?
24GB VRAM and 300 GB/s bandwidth rule out anything beyond small-model inference. A 70B-class model doesn't fit even quantized, and it has no NVLink for multi-card scale-up.
Is renting cheaper than buying?
Renting avoids the upfront hardware cost and lets you match spend to actual usage. A rented L4 at $0.649/hr only costs money while it's running, whereas buying ties up capital in hardware that keeps depreciating whether it's in use or not. Which is cheaper depends on how continuously you'd run it; short or bursty workloads usually favor renting. Out of capacity right now.
How is the L4 price calculated?
All prices are normalized to a per-GPU hourly rate using each offer's authoritative GPU count, which the raw price is divided by; some offers report price as already per-GPU. Offers whose price can't be safely normalized, or whose rate is an extreme outlier against the rest of the market, are excluded.
L4 price by region
Related guides
Other models in the same generation, then the rest of the GPU index.