NVIDIA Rubin GPU: Specs, Status and Release (2026)

Back
Team Aquanode

Team Aquanode

Sarthak Vaish

Updated OCTOBER 8, 2026Published OCTOBER 8, 2026

NVIDIA Rubin is the successor to Blackwell: a two-die GPU with up to 288 GB of HBM4 and, by NVIDIA's numbers, 50 PFLOPS of NVFP4 inference compute per GPU. As of October 8, 2026, NVIDIA says the platform is in full production and ramping, but it is sold through NVIDIA's sales channel and partner clouds, not as a GPU you can rent by the hour from a self-serve console today.

This guide covers what Rubin is, what NVIDIA has published about it, where its status stands as of October 2026, and what to run on in the meantime. It is part of our datacenter GPU series.

TL;DR

  • Rubin is NVIDIA's next datacenter GPU generation after Blackwell and Blackwell Ultra. NVIDIA lists 336 billion transistors, up to 288 GB of HBM4, and 50 PFLOPS of NVFP4 inference per GPU (NVIDIA's figures).
  • Status, with dates: NVIDIA said on January 5, 2026 that Rubin "is in full production" with partner products in the second half of 2026. Its August 26, 2026 earnings release says Vera Rubin is "ramping into full production with racks running at partners."
  • The naming changed: the rack once called NVL144 is now called Vera Rubin NVL72. The 72 counts GPU packages, the older 144 counted dies.
  • Aquanode does not rent Rubin today. The practical bridge is Blackwell Ultra (B300), B200 and H200, which run the same CUDA stack and the same low-precision formats.

Verdict: plan for Rubin as a 2027 capacity story for most teams, and spend 2026 getting your software ready for NVFP4 and large NVLink domains on Blackwell.

What Rubin is

Rubin is the GPU in NVIDIA's Vera Rubin platform, announced in detail at CES on January 5, 2026. The platform is built from six new chips: the Vera CPU, the Rubin GPU, the NVLink 6 switch, the ConnectX-9 SuperNIC, the BlueField-4 DPU and the Spectrum-6 Ethernet switch (NVIDIA newsroom, January 5, 2026). At GTC on March 16, 2026 NVIDIA added a seventh chip, the Groq 3 LPU, for low-latency decode.

Third-party coverage and roadmap sites sometimes use internal-looking names such as R100 or R200 for the GPU. NVIDIA's own pages simply say "Rubin GPU," so this guide does too.

Each Rubin GPU is a dual-die package, two reticle-sized dies on TSMC's 3 nm process, according to ServeTheHome's coverage of the CES launch. NVIDIA's technical blog gives 224 streaming multiprocessors and 336 billion transistors, against 208 billion for Blackwell.

Rubin specs against Blackwell and Hopper

The table uses NVIDIA's published numbers. Rubin figures are NVIDIA's claims for a product that is only now ramping, so treat them as vendor specifications, not measurements. Rows that NVIDIA's pages do not state are marked "not published."

SpecHopper H200Blackwell Ultra B300Rubin
Memory per GPU141 GB HBM3eUp to 288 GB HBM3eUp to 288 GB HBM4
Memory bandwidth per GPU4.8 TB/sNot published on the pages we read22 TB/s (technical blog) or 19.2 TB/s (NVL72 product page)
NVLink per GPU900 GB/s1.8 TB/s3.6 TB/s (technical blog) or 3 TB/s (NVL72 product page)
Headline low-precision compute3,958 TFLOPS FP8 with sparsity1.5x Blackwell's AI FLOPS (NVIDIA)50 PFLOPS NVFP4 inference, 35 PFLOPS NVFP4 training
TransistorsNot publishedNot published336 billion
Dies per GPU122

Two honest caveats. First, NVIDIA's own pages disagree on Rubin's memory bandwidth and NVLink figure: the January technical blog and the HGX page say 22 TB/s and 3.6 TB/s, while the Vera Rubin NVL72 product page lists 19.2 TB/s and 3 TB/s per GPU. Until NVIDIA reconciles them, quote both or quote neither. Second, "50 PFLOPS" is a sparse NVFP4 inference number from a Transformer Engine with adaptive compression. It is not comparable to an FP8 or BF16 figure, and it is not a prediction of your model's speedup. The NVFP4 training number (35 PFLOPS) is dense.

If the format names are new to you, read our NVFP4 vs MXFP4 explainer and the glossary entries for FP4 and HBM.

Architecture and form factors

The GPU. Beyond the two-die package, the headline change is memory: HBM4 replaces HBM3e, which lifts both capacity and bandwidth. For how the two memory generations differ, see HBM3e vs HBM4.

NVLink 6. Each GPU connects through the sixth-generation NVLink. In an NVL72 rack, 72 GPUs sit in a single all-to-all NVLink domain. NVIDIA's technical blog quotes 260 TB/s of aggregate switch bandwidth per rack, while the product page lists 216 TB/s; the same disagreement as above. See what NVLink is for the basics.

The Vera CPU. Vera has 88 custom Olympus cores with 176 threads, up to 1.5 TB of LPDDR5X at up to 1.2 TB/s, and a 1.8 TB/s NVLink-C2C link to the GPUs (NVIDIA technical blog, January 2026).

Form factors. NVIDIA lists several:

  • Vera Rubin NVL72: the rack-scale system with 72 Rubin GPUs and 36 Vera CPUs. The rack totals on NVIDIA's product page are 3,600 PFLOPS of NVFP4 inference, 2,520 PFLOPS of NVFP4 training, 20.7 TB of HBM4 and up to 54 TB of LPDDR5X. Our Vera Rubin NVL72 guide goes deeper.
  • HGX Rubin NVL8: an eight-GPU board for server makers, listed with 2 TB of HBM4 and 400 PFLOPS of NVFP4 inference.
  • DGX Rubin NVL8 and DGX Vera Rubin NVL72: NVIDIA's own systems built on those designs.
  • Vera Rubin NVL4: four Rubin GPUs and two Vera CPUs linked by NVLink-C2C, aimed at science workloads.

Status as of October 2026

The question people actually search is "can I get it." Here is the dated trail, all from NVIDIA unless noted:

DateStatementSource
January 5, 2026Rubin "is in full production, and Rubin-based products will be available from partners the second half of 2026." Cloud providers named for 2026 deployments.NVIDIA newsroom
March 16, 2026GTC 2026: seven chips in the platform, Vera Rubin NVL72 as the flagship rack.NVIDIA technical blog
August 26, 2026Earnings release: "Vera Rubin, now in full production." The platform "is ramping into full production with racks running at partners."NVIDIA Form 8-K
October 2026NVIDIA's Vera Rubin NVL72 page is labeled "Available Now," with a Contact Sales link and server makers shipping.NVIDIA product page

Read that plainly: production has started and racks are running at large cloud and infrastructure operators, but the purchase path is through NVIDIA and its partners. NVIDIA's releases give no public on-demand price and no general-availability date. Analyst estimates of 2026 Rubin volume exist, but they are estimates, not NVIDIA disclosures, and we do not repeat them here.

Aquanode does not rent Rubin. Our Vera Rubin page tracks the chip, and the live box below shows exactly what you can rent today.

What NVIDIA claims Rubin delivers

Everything in this section is NVIDIA's own claim from its January 5, 2026 announcement, not an independent benchmark:

  • Up to 10x lower cost per token than Blackwell for large mixture-of-experts inference.
  • 4x fewer GPUs to train mixture-of-experts models, compared with the predecessor platform (the summary names Blackwell, the body says predecessor).
  • Per ServeTheHome's reading of NVIDIA's CES material, about 5x Blackwell on NVFP4 inference per GPU and about 3.5x on training, with roughly 8x the inference performance per watt.

These are best-case figures for specific models and configurations. We have not measured them and neither has MLPerf for Rubin at the time of writing. For your workload, the useful question is whether your model is bound by memory bandwidth, memory capacity or interconnect, because those are the three things Rubin changes most. Our mixture-of-experts and KV cache glossary pages cover why those matter for inference.

Infrastructure needs

Rubin racks are liquid cooled. NVIDIA says Vera Rubin NVL72 systems "use warm-water, single-phase direct liquid cooling (DLC) with a 45-degree Celsius supply temperature," and describes the design as nearly doubling thermal performance in the same rack footprint compared with Blackwell. NVIDIA does not publish a per-GPU power figure for Rubin on the pages we reviewed, so we do not state one.

On the network side, each Rubin GPU gets 1.6 Tb/s of ConnectX-9 bandwidth, with Quantum-X800 InfiniBand or Spectrum-X Ethernet for scale-out. For the networking trade-off, see InfiniBand vs Ethernet for GPU clusters.

The software side is the good news. NVIDIA says software support is enabled with platform availability, and the programming model is the same CUDA stack. Work you do on NVFP4 and on large NVLink domains today carries forward.

When to choose Rubin, and when not to

  • Choose Rubin when you serve trillion-parameter mixture-of-experts models where per-token cost at scale dominates, you can site liquid-cooled racks or rent them from a cloud that does, and you can wait for volume.
  • Stay on Blackwell Ultra when you need capacity this quarter. B300 has up to 288 GB of HBM3e per GPU and runs NVFP4 today. See the B300 guide and B300 vs B200.
  • Stay on Hopper when your model fits in 141 GB, you serve in FP8 or BF16, and price per hour matters more than peak throughput. See the H200 page.

For the three-way comparison, read Rubin vs Blackwell vs Hopper.

Cost

We do not quote a Rubin rental price because there is no public on-demand price to quote. To estimate the economics of any GPU, take throughput from a cited benchmark for your model, convert it to tokens per GPU-hour, and multiply by the live hourly price in the box below. For Rubin there is no independent throughput figure yet, so any cost-per-token number you see for it today is the vendor's claim.

What to run today

Rubin is not something Aquanode rents. Aquanode manages and optimizes GPUs for training and inference workloads, and you can rent the GPUs in the box below on demand. The Blackwell and Hopper parts here run the same CUDA stack and the low-precision formats you would use on Rubin.

What's next

NVIDIA's roadmap puts a higher-density rack architecture called Kyber and a Rubin Ultra part in 2027, according to press coverage of GTC 2026; we could not confirm those dates or specifications on an NVIDIA page, so treat them as reported, not announced. The Rubin CPX context accelerator announced in 2025 has had a more complicated path, covered in our Rubin CPX guide.

FAQ

Is NVIDIA Rubin available to rent?

Not from Aquanode. NVIDIA says the platform is in full production with racks running at partners as of its August 26, 2026 earnings release, and its product page routes purchases through sales. Check the live box above for what Aquanode offers.

What is the difference between Rubin and Vera Rubin?

Rubin is the GPU. Vera is the CPU. Vera Rubin is the platform that pairs them with NVLink 6, ConnectX-9, BlueField-4 and Spectrum-6, and Vera Rubin NVL72 is the rack with 72 GPUs and 36 CPUs.

Is it NVL144 or NVL72?

NVIDIA's current materials say NVL72. The earlier NVL144 name counted dies (two per GPU); the new name counts GPU packages. They describe the same 72-GPU rack.

How much memory does Rubin have?

Up to 288 GB of HBM4 per GPU, according to NVIDIA. The rack total is 20.7 TB of HBM4.

How much faster is Rubin than Blackwell?

NVIDIA claims about 5x on NVFP4 inference per GPU and up to 10x lower cost per token on large mixture-of-experts models. These are vendor claims. Independent results for Rubin were not available when we wrote this.

Sources

#datacenter gpu#nvidia rubin#rubin gpu#hbm4#nvfp4#nvidia roadmap

Submit the job. Everything after that is ours.

Sign up in 60 seconds. Pay for the GPU minutes you actually use.

© 2026 Aquanode. All rights reserved.

All trademarks, logos and brand names are the property of their respective owners.