NVIDIA Rubin CPX: Specs and Current Status (2026)

Back
Team Aquanode

Team Aquanode

Sarthak Vaish

Updated OCTOBER 8, 2026Published OCTOBER 8, 2026

NVIDIA Rubin CPX is a GPU NVIDIA announced on September 9, 2025 for the context (prefill) phase of long-context inference: 128 GB of GDDR7, up to 30 PFLOPS of NVFP4 compute, and an expected availability of "the end of 2026." As of October 8, 2026, it is not shipping and NVIDIA has not given it a confirmed date: it was absent from NVIDIA's GTC 2026 roadmap in March, and a September 2026 analyst report says a redesigned version is slated for production in 2027, which NVIDIA has not confirmed.

This guide separates what NVIDIA announced, what it later said, and what is only reported. It is part of our datacenter GPU series.

TL;DR

  • Announced specs (NVIDIA, September 9, 2025): single monolithic die, 128 GB GDDR7, up to 30 PFLOPS NVFP4, 3x faster attention than GB300 NVL72 systems (NVIDIA's claim), four NVENC and four NVDEC video engines.
  • Announced availability: "expected to be available at the end of 2026." That date has not been met and NVIDIA has not restated it.
  • March 2026: Rubin CPX did not appear in the GTC 2026 keynote. NVIDIA instead showed the Groq 3 LPX rack for low-latency decode. Per Tom's Hardware, NVIDIA's Ian Buck said in a press Q&A that CPX had been pulled to focus on LPU-based decode this year.
  • September 2026 (unconfirmed): analyst Ming-Chi Kuo, as reported by RCR Wireless, says a redesigned CPX with HBM4 is headed for production in the first quarter of 2027.
  • Aquanode does not rent Rubin CPX and cannot offer it, a pre-order or a waitlist.

Verdict: do not plan a deployment around CPX. The idea behind it, splitting prefill from decode, is real and works on hardware you can use now.

What Rubin CPX was announced to do

LLM inference has two phases. The context phase (prefill) reads the whole prompt and builds the key-value cache, and it is compute-bound. The generation phase (decode) produces tokens one at a time and is mostly memory-bandwidth-bound. NVIDIA's pitch was that one GPU design is a compromise for both, so a chip optimized for prefill, with cheaper GDDR7 instead of HBM, would cut cost for long prompts such as large codebases and long video.

NVIDIA described Rubin CPX as a "new class of GPU designed for massive-context inference," meant for context windows beyond one million tokens, working alongside regular Rubin GPUs that handle decode. Read the glossary entry on the KV cache for why prefill output is large.

Announced specs

All figures below are from NVIDIA's September 9, 2025 press release unless noted. They are vendor figures for a product that has not shipped.

SpecRubin CPX (as announced)
Memory128 GB GDDR7
ComputeUp to 30 PFLOPS NVFP4
Attention3x faster than NVIDIA GB300 NVL72 systems (NVIDIA's claim)
Video4 NVENC and 4 NVDEC engines
DieSingle monolithic die
Rack configurationVera Rubin NVL144 CPX: 8 exaflops of AI compute, 100 TB of fast memory, 1.7 PB/s memory bandwidth in one rack
Rack vs GB300 NVL727.5x the AI performance (NVIDIA's claim)
Announced availabilityEnd of 2026

The single-die note comes from coverage of the announcement and the technical details; the rest are in the press release. The 30 PFLOPS number uses NVFP4, NVIDIA's four-bit format, and should be compared only with other NVFP4 figures. If the format is new to you, read NVFP4 vs MXFP4.

The press release also made an economics claim: $5 billion in token revenue for every $100 million invested in the CPX-based rack. That is a vendor revenue projection, not a measurement, and it depends on pricing assumptions NVIDIA does not publish.

One naming note. "Vera Rubin NVL144 CPX" is the 2025 rack name from the announcement. NVIDIA has since renamed its main Vera Rubin rack from NVL144 to NVL72 (counting GPU packages, not dies), so the 144 in the CPX rack name predates that change. See the Vera Rubin NVL72 guide.

How the disaggregated design was meant to work

Third-party summaries of NVIDIA's platform slide describe two layouts: a tray that pairs eight CPX GPUs with eight Rubin GPUs, and a CPX-only tray with no scale-up NVLink, where prefill output is handed to HBM-based Rubin nodes over the network (Glenn K. Lockwood's notes on the slides, April 1, 2026). In both, the flow is: prefill on CPX, ship the keys and values to the decode GPUs, generate tokens there.

This is the same idea as disaggregated serving in open inference engines, where prefill and decode run on separate GPU pools. You do not need CPX hardware to experiment with that pattern: you can run prefill and decode on different instances of GPUs you can rent now and measure the transfer cost yourself.

Status: what NVIDIA said, in order

DateWhat happenedSource
September 9, 2025Rubin CPX announced; "expected to be available at the end of 2026"NVIDIA newsroom
January 5, 2026Rubin platform launch at CES: six chips, no CPX among themNVIDIA newsroom
March 16, 2026GTC 2026 keynote: Groq 3 LPX rack shown for decode; CPX not mentionedTom's Hardware, NVIDIA technical blog
March 23, 2026Press Q&A transcript published: Ian Buck on shelving CPX and shipping LPU decode this yearTom's Hardware
September 2, 2026Analyst report of a redesign with HBM4 and production in Q1 2027; not confirmed by NVIDIARCR Wireless

What is sourced: NVIDIA's own January 5 platform list has six chips and no CPX, and its March 16 technical post on the Groq 3 LPX rack does not mention CPX. Tom's Hardware reported the omission from the keynote and published the Q&A, in which Buck is quoted saying CPX had been pulled so the company could focus on decoding with the LPU. The transcript is Tom's Hardware's own record of the session and notes it can occasionally be unclear; in it Buck says, "we've pulled CPX. It's still a good idea," and declines to rule out a 2026 CPX launch when asked, though the article itself only says the omission "may indicate" a shift to Groq 3 LPU.

What is only reported: RCR Wireless relays a post by supply-chain analyst Ming-Chi Kuo saying the revived part drops the 128 GB of GDDR7 for 168 GB of HBM4, rises to a 2,300 W power envelope, moves into standalone MGX racks, and starts production in Q1 2027. The article itself states that none of this is confirmed by NVIDIA. We repeat it for completeness and do not rely on it. Other outlets have described CPX as cancelled or as postponed toward a later generation; NVIDIA's own pages we reviewed neither cancel nor re-date it.

So the honest status is: announced, not shipping, no current NVIDIA date.

What replaced it in NVIDIA's story: Groq 3 LPX

NVIDIA's March 16, 2026 technical post describes the Groq 3 LPX rack as a low-latency inference accelerator for Vera Rubin, designed to run alongside Vera Rubin NVL72. Its published rack specs are 256 LPUs (Groq 3 LP30), 128 GB of on-chip SRAM, 40 PB/s of SRAM bandwidth, 640 TB/s of scale-up bandwidth and 315 PFLOPS of FP8. StorageReview reports NVIDIA saying the rack will be available in the second half of 2026, initially for model builders and service providers rather than broad OEM channels.

It targets decode, the opposite half of the problem CPX targeted. StorageReview's reading, which is the author's and not NVIDIA's, is that CPX's long-context concept has evolved into the LPX rack.

Performance: what is and is not published

NVIDIA's only CPX performance statements are the ones above: 30 PFLOPS NVFP4, 3x attention versus GB300 NVL72 systems, and rack-level totals. We found no MLPerf result and no independent benchmark for Rubin CPX. We do not estimate any, and we have no tokens-per-GPU-hour figure for it.

Infrastructure

The announced CPX racks were part of the Vera Rubin family and so would use the same MGX rack ecosystem and liquid cooling as Vera Rubin NVL72. NVIDIA published no CPX power figure on the pages we reviewed. The only power number in circulation, 2,300 W, comes from the unconfirmed analyst report above. The Vera Rubin cooling details NVIDIA has published are in the Rubin GPU guide.

When long-context work should use what

  • Prompts up to a few hundred thousand tokens, any model that fits: B300 has up to 288 GB of HBM3e per GPU, which leaves room for large KV caches. NVIDIA's Blackwell Ultra material cites 2x attention-layer acceleration over Blackwell.
  • Memory-capacity-bound serving on Hopper: H200 has 141 GB at 4.8 TB/s (NVIDIA H200 page).
  • Disaggregation experiments: run prefill and decode on separate pools today and measure. The mixture-of-experts glossary entry covers why decode batches behave differently.

Cost

There is no CPX price and no rental price. For the GPUs you can use, take a cited throughput for your model, convert to tokens per GPU-hour, and multiply by the live hourly price in the box below. We do not type hourly prices in this post.

What to run today

Rubin CPX is not something Aquanode rents, and there is no way to reserve it through us. Aquanode manages and optimizes GPUs for training and inference workloads, and you can rent the GPUs in the box below on demand. They are the best hardware available to prototype long-context serving now.

What's next

NVIDIA's roadmap pages we reviewed list Rubin and Vera Rubin NVL72 for the second half of 2026 and the LPX rack alongside it. For CPX, the next dated fact would have to come from NVIDIA. Until then, the Q1 2027 report is a lead, not a schedule. For the three-generation picture, read Rubin vs Blackwell vs Hopper.

FAQ

Is Rubin CPX available?

No. NVIDIA announced it for the end of 2026, but it has not shipped and NVIDIA has not given a current date. Aquanode does not rent it.

Was Rubin CPX cancelled?

NVIDIA has not published a cancellation. It left CPX out of its GTC 2026 roadmap, and Tom's Hardware reports Ian Buck saying it was pulled to focus on LPU decode. A September 2026 analyst report claims a redesign is coming in 2027, unconfirmed.

What memory does Rubin CPX use?

As announced, 128 GB of GDDR7 rather than HBM. The analyst report says a redesigned version would use 168 GB of HBM4; that is unconfirmed.

What is the difference between Rubin and Rubin CPX?

Rubin is the main GPU with HBM4, built for both phases of inference and for training. CPX was announced as a cheaper-memory GPU focused on the compute-bound context phase.

What is the Groq 3 LPX?

A rack of 256 Groq 3 LPUs with on-chip SRAM, designed by NVIDIA to run beside Vera Rubin NVL72 and accelerate low-latency decode.

Sources

#datacenter gpu#nvidia rubin#rubin cpx#long-context inference#prefill decode#groq 3 lpx

Submit the job. Everything after that is ours.

Sign up in 60 seconds. Pay for the GPU minutes you actually use.

© 2026 Aquanode. All rights reserved.

All trademarks, logos and brand names are the property of their respective owners.