NVIDIA Vera Rubin NVL72 is a liquid-cooled rack that joins 72 Rubin GPUs and 36 Vera CPUs into one NVLink domain, with 20.7 TB of HBM4 across the rack according to NVIDIA's product page. It is the same system that early coverage called Vera Rubin NVL144: the name changed from counting dies to counting GPU packages, and as of October 2026 NVIDIA says the platform is in full production and ramping at partners.
This guide covers the naming, the rack's published specs, cooling and power, its status with dates, and what to run today. It sits in our datacenter GPU series next to the Rubin GPU guide.
TL;DR
- Vera Rubin NVL72: 72 Rubin GPUs, 36 Vera CPUs, sixth-generation NVLink, ConnectX-9 SuperNICs and BlueField-4 DPUs in one rack (NVIDIA, January 5, 2026).
- Rack totals on NVIDIA's product page: 3,600 PFLOPS NVFP4 inference, 2,520 PFLOPS NVFP4 training, 20.7 TB HBM4, up to 54 TB LPDDR5X.
- NVL144 and NVL72 are the same rack. The 144 counted GPU dies (two per GPU); 72 counts packages. NVIDIA's current materials use NVL72.
- Status: "ramping into full production with racks running at partners" (NVIDIA Form 8-K, August 26, 2026). Aquanode does not rent it.
Verdict: NVL72 is the right mental model for what a Rubin deployment looks like. If you are planning for it, the work to do now is on the Blackwell racks that exist today.
NVL144 vs NVL72: the naming
When NVIDIA first described the platform in 2025, the rack was labeled Vera Rubin NVL144. At CES on January 5, 2026 it became Vera Rubin NVL72. ServeTheHome's coverage of the launch states that the rack was "previously known as NVL144," that each Rubin GPU is a dual-die package, and that the 72 counts GPU packages while the old 144 counted dies.
For anything you read or write, the safe rule is: NVL72 means 72 GPU packages in one NVLink domain, which is the same counting NVIDIA uses for GB200 NVL72 and GB300 NVL72. Older articles with "NVL144" and "3.6 exaflops" or "8 exaflops" are describing the same rack or the CPX variant, so check the date before comparing numbers.
There is one more name to keep apart. The "Vera Rubin NVL144 CPX" rack from 2025 paired Rubin GPUs with Rubin CPX accelerators. That is a separate configuration whose status is different; see the Rubin CPX guide.
Vera Rubin NVL72 specs
NVIDIA's Vera Rubin NVL72 product page lists the rack, the two-GPU Superchip and the single GPU. These are NVIDIA's published specifications for a product that is only now ramping, not measured results.
| Spec | NVL72 rack | Superchip (2 GPUs, 1 CPU) | Rubin GPU |
|---|---|---|---|
| GPUs | 72 | 2 | 1 |
| CPUs | 36 Vera | 1 Vera | none |
| NVFP4 inference (sparse) | 3,600 PFLOPS | 100 PFLOPS | 50 PFLOPS |
| NVFP4 training (dense) | 2,520 PFLOPS | 70 PFLOPS | 35 PFLOPS |
| FP8/FP6 training (dense) | 1,260 PFLOPS | 35 PFLOPS | 17.5 PFLOPS |
| HBM4 capacity | 20.7 TB | 576 GB | 288 GB |
| HBM4 bandwidth | 1,400 TB/s | 38.5 TB/s | 19.2 TB/s |
| LPDDR5X (CPU memory) | Up to 54 TB | Up to 1.5 TB | not applicable |
| NVLink 6 bandwidth | 216 TB/s | 6 TB/s | 3 TB/s |
A caveat you should know about: NVIDIA's January technical blog gives different bandwidth numbers for the same hardware, namely 22 TB/s of HBM4 bandwidth and 3.6 TB/s of NVLink per GPU, and 260 TB/s of NVLink switch bandwidth per rack. The product page (the table above) says 19.2 TB/s, 3 TB/s and 216 TB/s. We show the product-page values here and flag the discrepancy rather than pick a winner.
The Vera CPU is its own part: 88 custom Olympus cores, 176 threads, up to 1.5 TB of LPDDR5X at up to 1.2 TB/s, and a 1.8 TB/s NVLink-C2C link to the GPUs (NVIDIA technical blog, January 2026).
Architecture: how the rack is built
One NVLink domain. The 72 GPUs communicate over NVLink 6 in an all-to-all topology, so the rack behaves like one large accelerator for tensor-parallel and expert-parallel work. NVIDIA's blog describes each compute tray as 200 PFLOPS of NVFP4 with 2 TB of fast memory and 14.4 TB/s of NVLink 6 bandwidth. For the concept, see what NVLink is and the glossary entry on NVSwitch.
Networking out of the rack. Each Rubin GPU gets 1.6 Tb/s of ConnectX-9 SuperNIC bandwidth, and the rack scales out over Quantum-X800 InfiniBand or Spectrum-X Ethernet, with Spectrum-6 switches (NVIDIA technical blog). The trade-offs are in InfiniBand vs Ethernet for GPU clusters.
Systems built from it. NVIDIA says a DGX SuperPOD is built from eight DGX Vera Rubin NVL72 systems. NVIDIA also lists a smaller eight-GPU form, HGX Rubin NVL8, listed at 2 TB of HBM4 and 400 PFLOPS of NVFP4 inference, and the four-GPU Vera Rubin NVL4. For how HGX, DGX and NVL racks differ in general, read HGX vs DGX vs NVL72.
The extra racks. NVL72 is one of several rack types in the platform. NVIDIA's Rubin page also lists a Vera CPU rack (256 Vera CPUs per liquid-cooled rack), a Groq 3 LPX rack for low-latency decode, a BlueField-4 STX storage rack and a Spectrum-6 SPX networking rack.
Against GB200 and GB300 NVL72
The closest comparison is the Blackwell-generation rack. NVIDIA's GB300 NVL72 page lists 72 Blackwell Ultra GPUs, 36 Grace CPUs, 20 TB of HBM3e, 130 TB/s of NVLink bandwidth and 1,440 PFLOPS of FP4 with sparsity (the page footnotes also give 1,080 without sparsity).
| Rack | GPUs | GPU memory | NVLink bandwidth | FP4 / NVFP4 inference (sparse) |
|---|---|---|---|---|
| GB300 NVL72 | 72 Blackwell Ultra | 20 TB HBM3e | 130 TB/s | 1,440 PFLOPS |
| Vera Rubin NVL72 | 72 Rubin | 20.7 TB HBM4 | 216 TB/s (product page) | 3,600 PFLOPS |
Two things stand out. Memory capacity barely moves at the rack level (20 TB to 20.7 TB), so Rubin's gains come from bandwidth, interconnect and low-precision compute, not from holding bigger models in one domain. And the NVFP4 inference number is a sparse, format-specific peak, so the 2.5x ratio between those two cells is a ratio of peak specs, not a promise about your throughput. For the Blackwell side of this comparison, read GB300 NVL72 vs GB200 NVL72.
NVIDIA's own headline claims for the platform, from its January 5, 2026 announcement, are up to 10x lower cost per token than Blackwell for large mixture-of-experts inference, and 4x fewer GPUs to train mixture-of-experts models. Those are NVIDIA's claims. We have not measured them and we are not aware of independent Rubin results yet.
Power, cooling and facility needs
Vera Rubin NVL72 uses warm-water, single-phase direct liquid cooling with a 45 degrees Celsius supply temperature, and NVIDIA says it nearly doubles thermal performance in the same rack footprint compared with Blackwell. It also says the rack has about 6x more local energy buffering than Blackwell Ultra, and that its DSX reference design can enable up to 30% more GPU capacity in the same power envelope (NVIDIA technical blog, January 2026).
NVIDIA does not publish a rack power figure on the pages we reviewed, so we do not state one. Treat any kilowatt number you see for Rubin racks as a third-party estimate. Our practical advice: if you cannot host liquid-cooled racks, you will consume this hardware through someone who can.
Status as of October 8, 2026
| Date | Statement | Source |
|---|---|---|
| January 5, 2026 | Rubin "is in full production"; partner products in the second half of 2026; first clouds deploying Vera Rubin instances in 2026 include AWS, Google Cloud, Microsoft and OCI | NVIDIA newsroom |
| March 16, 2026 | GTC 2026: Vera Rubin NVL72 is the flagship of a seven-chip, five-rack platform; LPX rack to ship in the second half of 2026 | NVIDIA technical blog, StorageReview |
| August 26, 2026 | Vera Rubin "ramping into full production with racks running at partners" | NVIDIA Form 8-K |
| October 2026 | Product page labeled "Available Now" with Contact Sales; Taiwanese server makers shipping systems | NVIDIA product page |
The plain reading: racks exist and are running at large operators, and NVIDIA sells through its sales channel and partners. NVIDIA does not publish a self-serve hourly price. Aquanode does not rent Vera Rubin, and this post is not a way to reserve it. Our Vera Rubin page follows the chip.
When to choose a rack-scale system
Rack-scale NVLink domains pay off for workloads that need all 72 GPUs to act as one: very large mixture-of-experts serving with expert parallelism, long-context inference that shards the KV cache across GPUs, and training where the tensor-parallel group is bigger than eight. Our glossary entries on tensor parallelism, mixture of experts and KV cache explain why.
If your model fits in an 8-GPU node, rack-scale gains are smaller than they look in headline numbers, and a B300 or B200 node is the practical choice today.
Cost
There is no public hourly price for Vera Rubin NVL72, and we do not type one for any GPU in this post. To estimate cost on hardware you can rent, take a cited tokens-per-second figure for your model, convert it to tokens per GPU-hour, and multiply by the live hourly price in the box below.
What to run today
Vera Rubin NVL72 is not something Aquanode rents. Aquanode manages and optimizes GPUs for training and inference workloads, and you can rent the GPUs in the box below on demand. Blackwell and Blackwell Ultra are the closest hardware you can use now to prepare for NVLink-domain and NVFP4 work.
What's next
NVIDIA's own pages mention the platform and its LPX rack for the second half of 2026. Press coverage of GTC 2026 describes a next rack architecture, Kyber, with 144 GPUs per rack arriving with Vera Rubin Ultra in 2027; we could not confirm those details on an NVIDIA page, so treat them as reported, not announced.
FAQ
Is Vera Rubin NVL72 the same as NVL144?
Yes. NVL144 was the earlier name, counting two dies per GPU. NVL72 counts GPU packages. NVIDIA's current materials use NVL72.
How many GPUs and CPUs are in a Vera Rubin NVL72 rack?
72 Rubin GPUs and 36 Vera CPUs, connected with sixth-generation NVLink.
How much memory does the rack have?
20.7 TB of HBM4 on the GPUs and up to 54 TB of LPDDR5X on the CPUs, per NVIDIA's product page.
Is Vera Rubin NVL72 shipping?
NVIDIA's August 26, 2026 release says it is ramping into full production with racks running at partners, and its product page says systems are shipping from Taiwanese server makers. It is sold through NVIDIA and partners.
Does it need liquid cooling?
Yes. NVIDIA describes warm-water direct liquid cooling with a 45 degrees Celsius supply temperature.
Sources
- NVIDIA Vera Rubin NVL72 product page: https://www.nvidia.com/en-us/data-center/vera-rubin-nvl72/
- NVIDIA newsroom, Rubin platform announcement, January 5, 2026: https://nvidianews.nvidia.com/news/rubin-platform-ai-supercomputer
- NVIDIA technical blog, Inside the NVIDIA Rubin Platform: https://developer.nvidia.com/blog/inside-the-nvidia-rubin-platform-six-new-chips-one-ai-supercomputer/
- NVIDIA Rubin platform page: https://www.nvidia.com/en-us/data-center/technologies/rubin/
- NVIDIA HGX page: https://www.nvidia.com/en-us/data-center/hgx/
- NVIDIA GB300 NVL72 page: https://www.nvidia.com/en-us/data-center/gb300-nvl72/
- NVIDIA second-quarter fiscal 2027 results (Form 8-K), August 26, 2026: https://www.sec.gov/Archives/edgar/data/0001045810/000104581026000073/q2fy27pr.htm
- ServeTheHome, NVIDIA launches Rubin at CES 2026: https://www.servethehome.com/nvidia-launches-next-generation-rubin-ai-compute-platform-at-ces-2026/
- StorageReview, GTC 2026 coverage: https://www.storagereview.com/news/nvidia-gtc-2026-rubin-gpus-groq-lpus-vera-cpus-and-what-nvidia-is-building-for-trillion-parameter-inference