NVIDIA RTX 4060 GPU: Specs, VRAM, Price & Benchmarks (2026)
The RTX 4060 is a real GPU (launched 2023), but no provider on Aquanode is listing it for rental right now, so there is no live hourly price to quote. Availability changes as providers add and retire hardware; the specs, VRAM and price context below still apply.
How much VRAM does the RTX 4060 have?
The RTX 4060 has 8GB GDDR6, with 272 GB/s of derived memory bandwidth (bus width × transfer rate; not a vendor-stated figure).
RTX 4060 VRAM calculator: check which models fit in its memory at each precision.
All models that fit in 8 GB: the open models whose weights and overhead fit, at native, FP8 and INT4 precision.
What can the RTX 4060 run?
Popular open models from small to frontier scale, with the memory each needs and how many RTX 4060 cards (8GB GDDR6 each) that takes.
| Model | As published | FP8 | INT4 |
|---|---|---|---|
| Qwen/Qwen3-8B 8.2B | BF16: ~18.3 GB, 3 GPUs | FP8: ~9.2 GB, 2 GPUs | INT4: ~4.6 GB, 1 GPU |
| Qwen/Qwen2.5-14B-Instruct 14.8B | BF16: ~33 GB, 5 GPUs | FP8: ~16.5 GB, 3 GPUs | INT4: ~8.3 GB, 2 GPUs |
| Qwen/Qwen3-32B 32.8B | BF16: ~73.2 GB, 10 GPUs | FP8: ~36.6 GB, 5 GPUs | INT4: ~18.3 GB, 3 GPUs |
| Qwen/Qwen-72B 72.3B | BF16: ~162 GB, 21 GPUs | FP8: ~80.8 GB, 11 GPUs | INT4: ~40.4 GB, 6 GPUs |
| MiniMaxAI/MiniMax-M2.7 228.7B | FP8: ~256 GB, 32 GPUs | – | INT4: ~128 GB, 16 GPUs |
| deepseek-ai/DeepSeek-R1 684.5B | FP8: ~765 GB, 96 GPUs | – | INT4: ~383 GB, 48 GPUs |
Estimates: weights at the stated precision plus a flat 20% for KV cache and overhead, at a moderate context length. A dash means the precision is not offered for that model (it is already published at that size). INT4 needs a published quantized checkpoint. Open any model for a per-GPU breakdown, or use the RTX 4060 VRAM calculator.
RTX 4060 specs
| Architecture | launched 2023 |
| VRAM | 8GB GDDR6 |
| Memory bandwidth | 272 GB/s |
| TDP | 115W |
| Form factor | PCIe |
Specs sourced from the vendor's public product page. See the source.
GPU Glossary: What is VRAM?, Tensor Cores, CUDA Cores, TFLOPS
Good for
A 115W Ada card with FP8-capable tensor cores: 7-8B models at INT4 fit and run at very low hourly cost.
Not good for
8GB on a 128-bit bus means 272 GB/s, the slowest decode of the 8GB cards here, and nothing above ~13B fits even at INT4. No NVLink, no ECC.
Related guides
Other models in the same generation. The full list is in the GPU index.
Get notified when the price drops
GPU supply moves hourly. Tell us what you're waiting for and we'll email you when a matching offer appears across any provider we track.