Rubin vs Blackwell vs Hopper is a comparison of three NVIDIA datacenter generations at three different points in their life. Hopper (H100, H200) is mature and widely rentable, Blackwell (B200, B300, GB200, GB300) is the current volume generation, and Rubin is the next one, which NVIDIA says is in full production and ramping as of its August 26, 2026 earnings release but is not something Aquanode rents.
This post lines up the published specifications, labels which numbers are vendor claims, and says which generation fits which job. It belongs to our datacenter GPU series.
TL;DR
- Memory per GPU: H200 141 GB HBM3e, B300 up to 288 GB HBM3e, Rubin up to 288 GB HBM4 (NVIDIA pages). Rubin's gain is bandwidth and compute, not capacity over B300.
- Interconnect per GPU: Hopper NVLink 900 GB/s, Blackwell 1.8 TB/s, Rubin 3 to 3.6 TB/s depending on which NVIDIA page you read.
- Low precision is the real change: Hopper tops out at FP8 for its headline number, Blackwell adds FP4 hardware, and Rubin quotes 50 PFLOPS of NVFP4 inference (sparse) per GPU.
- Today you can rent Hopper and Blackwell. Rubin is announced and ramping, sold through NVIDIA and partners.
Verdict: pick by what you can run this month. Hopper for fits-in-141 GB, price-sensitive work; Blackwell for large models and FP4; Rubin as a planning target.
Side-by-side specs
All numbers are from NVIDIA's own pages, each linked in Sources. "Not published" means NVIDIA's page we read does not state it. Rubin numbers are for a product only now ramping.
| Spec | Hopper H100 | Hopper H200 | Blackwell B300 (HGX) | Rubin |
|---|---|---|---|---|
| Memory per GPU | 80 GB | 141 GB HBM3e | Up to 288 GB HBM3e (about 2.1 TB per 8 GPUs) | Up to 288 GB HBM4 |
| Memory bandwidth | 3.35 TB/s | 4.8 TB/s | Not published on the pages we read | 22 TB/s or 19.2 TB/s |
| NVLink per GPU | 900 GB/s | 900 GB/s | 1.8 TB/s | 3.6 TB/s or 3 TB/s |
| Headline FP8 | 3,958 TFLOPS (sparse) | 3,958 TFLOPS (sparse, SXM) | 72 PFLOPS per 8 GPUs (sparse) | 130 PFLOPS FP8/FP6 training per 8 GPUs (dense) |
| Headline FP4 | none | none | 144 PFLOPS per 8 GPUs sparse, 108 dense | 50 PFLOPS NVFP4 inference per GPU (sparse), 35 training (dense) |
| Max TDP | Up to 700 W | Up to 700 W (SXM) | Not published here | Not published |
Notes on reading this table:
- Rubin has two values for bandwidth and NVLink because NVIDIA's own pages disagree: the January 2026 technical blog and HGX page say 22 TB/s and 3.6 TB/s; the Vera Rubin NVL72 product page says 19.2 TB/s and 3 TB/s per GPU.
- Per-8-GPU and per-GPU figures are not interchangeable. Eight B300 GPUs at 144 PFLOPS sparse FP4 is 18 PFLOPS each; the Rubin HGX NVL8 board is listed at 400 PFLOPS NVFP4 inference, which is 50 per GPU.
- Sparse numbers assume structured sparsity your model may not use. Compare dense to dense when you can.
What changed from Hopper to Blackwell
Memory. H200 raised Hopper to 141 GB of HBM3e at 4.8 TB/s. Blackwell Ultra goes to up to 288 GB of HBM3e per GPU (NVIDIA Blackwell Ultra technical blog). That moves a 70-billion-parameter-class model from needing tensor parallelism to fitting with room for KV cache; see the KV cache glossary.
Interconnect. NVLink doubles from 900 GB/s to 1.8 TB/s per GPU, and the NVL72 racks extend one NVLink domain to 72 GPUs. See what NVLink is.
Precision. Blackwell introduced hardware for four-bit formats, and Blackwell Ultra's Tensor Cores deliver 1.5x more AI compute FLOPS than Blackwell with 2x attention-layer acceleration (NVIDIA's claims). The format story is in NVFP4 vs MXFP4, and the Hopper side is in the Transformer Engine and FP8 guide.
Rack scale. GB300 NVL72 combines 72 Blackwell Ultra GPUs and 36 Grace CPUs with 20 TB of HBM3e and 130 TB/s of NVLink bandwidth (NVIDIA GB300 NVL72 page). NVIDIA claims, for GB300 NVL72 against an HGX H100 baseline, 10x tokens per second per user and 5x tokens per second per megawatt. Those are the vendor's claims for specific configurations.
Related reading: B300 vs B200, the B200 guide, and H200 vs B200 vs GB200.
What changes from Blackwell to Rubin
Memory. Same top capacity per GPU (288 GB), new technology: HBM4 instead of HBM3e. NVIDIA cites 22 TB/s of bandwidth on its technical blog, roughly 2.8x Blackwell according to ServeTheHome's reading of the CES material. See HBM3e vs HBM4.
Compute. NVIDIA lists 336 billion transistors against 208 billion for Blackwell, and 50 PFLOPS NVFP4 inference per GPU, about 5x Blackwell by ServeTheHome's account of NVIDIA's figures. These are peak vendor figures in a specific format with a new Transformer Engine; no independent Rubin benchmark was available when we wrote this.
Interconnect. NVLink 6 roughly doubles per-GPU bandwidth again, and the Vera Rubin NVL72 rack lists 216 TB/s of NVLink bandwidth against 130 TB/s for GB300 NVL72.
The CPU. The Vera CPU has 88 Olympus cores and a 1.8 TB/s NVLink-C2C link to the GPUs, replacing Grace.
Cooling. Vera Rubin NVL72 uses warm-water direct liquid cooling with a 45 degrees Celsius supply temperature (NVIDIA technical blog).
NVIDIA's headline claims for Rubin, from its January 5, 2026 announcement: up to 10x lower cost per token than Blackwell for large mixture-of-experts inference, and 4x fewer GPUs to train mixture-of-experts models. We have not verified either. For depth, see the Rubin GPU guide and the Vera Rubin NVL72 guide.
Status of each generation, October 2026
| Generation | Status | Source |
|---|---|---|
| Hopper (H100, H200) | Shipping for years; rentable | NVIDIA product pages |
| Blackwell (B200, B300, GB200, GB300) | Current volume generation; rentable | NVIDIA product pages |
| Rubin / Vera Rubin NVL72 | "Ramping into full production with racks running at partners" (August 26, 2026); product page lists Contact Sales | NVIDIA Form 8-K, product page |
| Rubin CPX | Announced for end of 2026; not shipping, absent from the GTC 2026 roadmap | See the Rubin CPX guide |
Aquanode does not rent Rubin or Rubin CPX. Our GPU pages show what is rentable, and the box below is live.
Which generation fits which workload
- Models that fit in 80 to 141 GB, FP8 or BF16, cost-sensitive: Hopper. See H100 and H200, and for SXM versus PCIe versus NVL variants, H100 and H200 form factors.
- Large models, long context, FP4 serving, fast training: Blackwell. B200 for the established path, B300 for the most memory per GPU. See B200 and B300.
- Rack-scale expert-parallel serving at very large scale, with liquid-cooled facilities: Rubin NVL72 when your cloud offers it. Until then, GB200 and GB300 racks are the available stand-in.
- Migration planning: write your serving stack to be format-agnostic. FP8 on Hopper, NVFP4 on Blackwell and Rubin is the direction every generation's headline number points.
For head-to-head GPU pages, see B200 vs H200 and B200 vs H100.
Cost
We do not quote Rubin prices because none are public, and we do not type hourly prices anywhere in this post. To compare generations on cost, take a published throughput for your model on each GPU, convert to tokens (or images) per GPU-hour, and multiply by the live hourly price in the box below. A generation that is 2x faster is cheaper per token only if it costs less than 2x per hour.
What to run today
Rubin is not something Aquanode rents. Aquanode manages and optimizes GPUs for training and inference workloads, and you can rent the GPUs in the box below on demand: two Blackwell parts and one Hopper part, the generations you can actually deploy now.
What's next
After Rubin, press coverage of GTC 2026 describes Vera Rubin Ultra and a 144-GPU Kyber rack in 2027; we could not confirm those on an NVIDIA page, so treat them as reported. The only dated Rubin-family facts we rely on are the ones in the status table above.
FAQ
Is Rubin better than Blackwell?
On NVIDIA's published peaks, yes: 50 PFLOPS NVFP4 inference per GPU against roughly a fifth of that for Blackwell by NVIDIA's own ratio, plus HBM4 and NVLink 6. Whether it is better for your workload depends on whether you are bound by compute, memory bandwidth or interconnect, and independent Rubin results were not available when we wrote this.
Is Rubin available yet?
NVIDIA says it is ramping into full production with racks at partners, and sells through its sales channel. Aquanode does not rent it.
Is Hopper obsolete?
No. It is mature, widely available and sufficient for any model that fits its memory. The decision is cost per token for your model, not the calendar.
How much memory does each have?
H100 80 GB, H200 141 GB, B300 up to 288 GB, Rubin up to 288 GB, per NVIDIA's pages.
What is the difference between NVL72 and NVL144?
The same Vera Rubin rack. NVL144 counted dies; NVL72 counts GPU packages.
Sources
- NVIDIA newsroom, Rubin platform, January 5, 2026: https://nvidianews.nvidia.com/news/rubin-platform-ai-supercomputer
- NVIDIA technical blog, Inside the NVIDIA Rubin Platform: https://developer.nvidia.com/blog/inside-the-nvidia-rubin-platform-six-new-chips-one-ai-supercomputer/
- NVIDIA Vera Rubin NVL72 product page: https://www.nvidia.com/en-us/data-center/vera-rubin-nvl72/
- NVIDIA HGX page: https://www.nvidia.com/en-us/data-center/hgx/
- NVIDIA GB300 NVL72 page: https://www.nvidia.com/en-us/data-center/gb300-nvl72/
- NVIDIA Blackwell Ultra technical blog: https://developer.nvidia.com/blog/nvidia-blackwell-ultra-for-the-era-of-ai-reasoning/
- NVIDIA H100 page: https://www.nvidia.com/en-us/data-center/h100/
- NVIDIA H200 page: https://www.nvidia.com/en-us/data-center/h200/
- NVIDIA second-quarter fiscal 2027 results (Form 8-K), August 26, 2026: https://www.sec.gov/Archives/edgar/data/0001045810/000104581026000073/q2fy27pr.htm
- ServeTheHome, NVIDIA launches Rubin at CES 2026: https://www.servethehome.com/nvidia-launches-next-generation-rubin-ai-compute-platform-at-ces-2026/
- StorageReview, GTC 2026 coverage: https://www.storagereview.com/news/nvidia-gtc-2026-rubin-gpus-groq-lpus-vera-cpus-and-what-nvidia-is-building-for-trillion-parameter-inference