NVIDIA Vera Rubin GPU: Specs, VRAM, Price & Benchmarks (2026)
The Vera Rubin is a real GPU (launched 2026), but no provider on Aquanode is listing it for rental right now, so there is no live hourly price to quote. Availability changes as providers add and retire hardware; the specs, VRAM and price context below still apply.
How much VRAM does the Vera Rubin have?
The Vera Rubin has 288GB HBM4 per GPU, with 22 TB/s of peak memory bandwidth.
What can the Vera Rubin run?
Popular open models from small to frontier scale, with the memory each needs and how many Vera Rubin cards (288GB HBM4 per GPU each) that takes.
| Model | As published | FP8 | INT4 |
|---|---|---|---|
| Qwen/Qwen3-8B 8.2B | BF16: not supported | FP8: not supported | INT4: not supported |
| Qwen/Qwen2.5-14B-Instruct 14.8B | BF16: not supported | FP8: not supported | INT4: not supported |
| Qwen/Qwen3-32B 32.8B | BF16: not supported | FP8: not supported | INT4: not supported |
| Qwen/Qwen-72B 72.3B | BF16: not supported | FP8: not supported | INT4: not supported |
| MiniMaxAI/MiniMax-M2.7 228.7B | FP8: not supported | – | INT4: not supported |
| deepseek-ai/DeepSeek-R1 684.5B | FP8: not supported | – | INT4: not supported |
Estimates: weights at the stated precision plus a flat 20% for KV cache and overhead, at a moderate context length. A dash means the precision is not offered for that model (it is already published at that size). INT4 needs a published quantized checkpoint. Open any model for a per-GPU breakdown.
Vera Rubin specs
| Architecture | launched 2026 |
| VRAM | 288GB HBM4 per GPU |
| Memory bandwidth | 22 TB/s |
| Interconnect | NVLink 6 (platform-disclosed; per-GPU bandwidth not yet published) |
| Form factor | Superchip module (rack-scale; deployed in the Vera Rubin NVL72 rack) |
Specs sourced from the vendor's public product page. See the source.
GPU Glossary: What is VRAM?, HBM, Tensor Cores, CUDA Cores, TFLOPS, NVLink vs PCIe
Good for
NVIDIA's announced successor to Blackwell: 288GB of HBM4 per GPU at 22 TB/s (2.8x Blackwell's HBM3e bandwidth), rated at 50 PFLOPS NVFP4 inference and 35 PFLOPS NVFP4 training per GPU. A full Vera Rubin NVL72 rack pairs 72 Rubin GPUs with 36 Vera CPUs for 3,600 PFLOPS of NVFP4 inference per rack.
Not good for
Not shippable to customers yet at the time of writing: NVIDIA describes it as entering full production in January 2026 and shipping to hyperscalers and neoclouds through H2 2026. No GPU marketplace, including Aquanode, lists per-GPU Vera Rubin rental today, and no rent CTA is shown on this page for that reason.
Related guides
Other models in the same generation. The full list is in the GPU index.
Get notified when the price drops
GPU supply moves hourly. Tell us what you're waiting for and we'll email you when a matching offer appears across any provider we track.