NVIDIA A16 GPU: Specs, VRAM, Price & Benchmarks (2026)
The A16 is a real GPU (launched 2021), but no provider on Aquanode is listing it for rental right now, so there is no live hourly price to quote. Availability changes as providers add and retire hardware; the specs, VRAM and price context below still apply.
How much VRAM does the A16 have?
The A16 has 16GB GDDR6 with ECC, with 200 GB/s of peak memory bandwidth.
A16 VRAM calculator: check which models fit in its memory at each precision.
All models that fit in 16 GB: the open models whose weights and overhead fit, at native, FP8 and INT4 precision.
What can the A16 run?
Popular open models from small to frontier scale, with the memory each needs and how many A16 cards (16GB GDDR6 with ECC each) that takes.
| Model | As published | FP8 | INT4 |
|---|---|---|---|
| Qwen/Qwen3-8B 8.2B | BF16: ~18.3 GB, 2 GPUs | FP8: not supported | INT4: ~4.6 GB, 1 GPU |
| Qwen/Qwen2.5-14B-Instruct 14.8B | BF16: ~33 GB, 3 GPUs | FP8: not supported | INT4: ~8.3 GB, 1 GPU |
| Qwen/Qwen3-32B 32.8B | BF16: ~73.2 GB, 5 GPUs | FP8: not supported | INT4: ~18.3 GB, 2 GPUs |
| Qwen/Qwen-72B 72.3B | BF16: ~162 GB, 11 GPUs | FP8: not supported | INT4: ~40.4 GB, 3 GPUs |
| MiniMaxAI/MiniMax-M2.7 228.7B | FP8: not supported | – | INT4: ~128 GB, 8 GPUs |
| deepseek-ai/DeepSeek-R1 684.5B | FP8: not supported | – | INT4: ~383 GB, 24 GPUs |
Estimates: weights at the stated precision plus a flat 20% for KV cache and overhead, at a moderate context length. A dash means the precision is not offered for that model (it is already published at that size). INT4 needs a published quantized checkpoint. Open any model for a per-GPU breakdown, or use the A16 VRAM calculator.
A16 specs
| Architecture | launched 2021 |
| VRAM | 16GB GDDR6 with ECC |
| Memory bandwidth | 200 GB/s |
| TDP | 250W (board total, shared across the 4 GPUs) |
| Form factor | PCIe, full height/full length, dual-slot (4-GPU board) |
Specs sourced from the vendor's public product page. See the source.
GPU Glossary: What is VRAM?, Tensor Cores, CUDA Cores, TFLOPS
Good for
A virtualization-focused Ampere card built for dense VDI, not AI throughput. Each of the 4 GPUs on the board is individually rentable here at 16GB GDDR6, a low-cost fit for a single 7-8B model in BF16 or a quantized larger one.
Not good for
200 GB/s of GDDR6 bandwidth per GPU is far below any HBM data-center card, no FP8 tensor cores (Ampere generation), and no NVLink between dies, so it never scales past one GPU's 16GB for a single job.
Related guides
Other models in the same generation. The full list is in the GPU index.
Get notified when the price drops
GPU supply moves hourly. Tell us what you're waiting for and we'll email you when a matching offer appears across any provider we track.