NVIDIA Groq 3 LPU and LPX Rack: Specs, Status (2026)

Back
Team Aquanode

Team Aquanode

Sarthak Vaish

Updated OCTOBER 8, 2026Published OCTOBER 8, 2026

The NVIDIA Groq 3 LPU is a language processing unit that NVIDIA brought into its Vera Rubin platform after a December 2025 licensing deal with Groq, and the Groq 3 LPX is the rack that holds 256 of them: 128 GB of on-chip SRAM, 40 PB/s of SRAM bandwidth and 315 PFLOPS of FP8, per NVIDIA's technical blog. It is a decode accelerator that sits beside Vera Rubin NVL72 to make token generation faster, not a GPU replacement, and NVIDIA's pages we read give no public shipping date or price.

This guide covers what an LPU is, the Groq deal, NVIDIA's published specs, how work is split with Rubin GPUs, and status as of October 8, 2026. It is part of our datacenter GPU series.

TL;DR

  • What it is: a rack of 256 Groq 3 LP30 LPUs, built on SRAM instead of HBM, designed to run alongside Vera Rubin NVL72 for low-latency decode (NVIDIA technical blog, March 16, 2026).
  • Where it came from: on December 24, 2025 Groq announced a non-exclusive licensing agreement with NVIDIA for its inference technology; founder Jonathan Ross, president Sunny Madra and other team members joined NVIDIA, and Groq continues as a separate company.
  • Specs: 256 LPUs, 128 GB SRAM, 40 PB/s SRAM bandwidth, 640 TB/s scale-up bandwidth, 315 PFLOPS FP8 per rack.
  • Status: NVIDIA's GTC 2026 announcement says LPX will be available in the second half of 2026 (NVIDIA newsroom); StorageReview adds an initial focus on model builders and service providers. Aquanode does not rent it.

Verdict: LPX targets one narrow job, very fast per-user token generation at trillion-parameter scale. It matters for how NVIDIA frames inference, but it is not hardware most teams will touch in 2026.

What an LPU is

A GPU keeps model weights and the key-value cache in HBM, which is large but comparatively distant from the compute units. An LPU, the processor Groq designed, keeps weights in on-chip SRAM, which is far faster but far smaller. NVIDIA's numbers show the trade: each Groq 3 LPU has 500 MB of SRAM with 150 TB/s of on-chip memory bandwidth, compared with the 288 GB of HBM4 on a Rubin GPU.

Because 500 MB cannot hold a large model, the design spreads one model across many chips connected by direct links. Each LPU has 96 chip-to-chip links at 112 Gbps and 2.5 TB/s of aggregate bidirectional I/O (NVIDIA technical blog). That makes the rack, not the chip, the unit of deployment.

If the memory terms are new, our glossary entries on HBM and the KV cache cover them, and HBM3e vs HBM4 compares the GPU side.

The Groq deal, with dates

DateEventSource
December 24, 2025Groq announces a "non-exclusive licensing agreement" with NVIDIA for its inference technology. Jonathan Ross, Sunny Madra and other team members join NVIDIA. Groq stays a separate company with Simon Edwards as CEO, and its cloud service continues.Groq announcement
December 2025CNBC reports about $20 billion in cash, per the CEO of Groq's lead investor; NVIDIA's CFO declined comment on the transaction.Press reports
March 16, 2026GTC 2026: Groq 3 LPU and LPX rack shown as part of the Vera Rubin platformNVIDIA technical blog

Two precision points. First, the structure is a license plus hiring, not an acquisition of the company, which is how Groq and NVIDIA both describe it. Second, the $20 billion figure comes from media reports, not from either company, and other outlets report slightly different totals. We cite it as reported.

The Groq 3 LPU is the first chip to come out of that arrangement, which means it is a first-generation product on NVIDIA's platform with no production track record yet.

Published specs

All figures are from NVIDIA's technical blog post "Inside NVIDIA Groq 3 LPX" (March 16, 2026) and NVIDIA's LPX product page, and are vendor specifications for a product not yet widely deployed.

SpecPer rack (256 LPUs)Per LPU
AI inference compute315 PFLOPS FP8not published per chip on pages we read
On-chip SRAM128 GB500 MB
SRAM bandwidth40 PB/s150 TB/s
Scale-up bandwidth640 TB/s2.5 TB/s aggregate bidirectional I/O
Attached DRAM12 TB DDR5 (product page)not applicable
Form factor32 liquid-cooled 1U compute trays in an MGX ETL rack8 LPUs per tray

Each tray carries 8 LPUs, 4 GB of SRAM, 1.2 PB/s of SRAM bandwidth and 9.6 PFLOPS of FP8, with up to 256 GB of DRAM via fabric expansion logic and up to 128 GB via the host CPU.

Early coverage quoted different SRAM bandwidth and scale-up figures; the NVIDIA blog and product page now agree on 40 PB/s and 640 TB/s, and those are the numbers above.

How the work is split with Rubin GPUs

NVIDIA's product page says LPX pairs Rubin GPUs, which use HBM, with LPUs, which use SRAM, and that the two jointly compute every layer for each output token. It describes the links between LPX and Vera Rubin NVL72 as reducing latency to near zero and says LPX is built "to pair with NVIDIA Vera Rubin NVL72" for low-latency, large-context agentic work. The product page focuses on decode and does not describe a separate prefill stage.

Third-party descriptions differ on the exact division. One reading is that GPUs keep the compute-heavy prefill and the attention part of decode, while the LPUs run the latency-sensitive feed-forward and mixture-of-experts work. That is an outside reading, not an NVIDIA statement, so treat the layer-by-layer split as not fully published. For why MoE decode is hard to make fast, see the glossary entry on mixture of experts.

The concept it replaced in NVIDIA's slides is Rubin CPX, a prefill-oriented accelerator announced in 2025 and left out of the GTC 2026 roadmap. LPX is the opposite half of the same prefill-versus-decode idea. Our Rubin CPX guide has the dated trail.

Performance claims (NVIDIA's, projected)

NVIDIA's LPX page marks its claims as projected and subject to change:

  • Up to 35x higher throughput per megawatt for trillion-parameter models, relative to the comparison NVIDIA describes (the page says throughput per megawatt).
  • Agentic systems consume up to 15x more tokens than traditional AI applications.
  • AI factories could unlock up to 10x more revenue per watt, based on projected throughput per gigawatt and a tiered pricing model.
  • The benchmark tiers named are Qwen-3 235B with 32K cached tokens, Kimi K2.5 1T with 128K, and a 2T-parameter MoE model at 128K and 400K.

These are projections from the vendor. We found no MLPerf result and no independent measurement for LPX, and we do not estimate any. For the other SRAM-centric approach in this market, wafer-scale chips, read Cerebras vs NVIDIA.

Infrastructure

LPX uses NVIDIA's MGX ETL rack architecture and liquid-cooled trays, and it is designed to be deployed beside Vera Rubin NVL72 racks, whose own cooling is warm-water direct liquid cooling at a 45 degrees Celsius supply temperature (NVIDIA technical blog). NVIDIA does not publish an LPX rack power figure on the pages we read, so we do not state one. The networking and rack details for the Rubin side are in the Rubin GPU guide.

Status as of October 8, 2026

DateStatementSource
March 16, 2026LPX rack shown at GTC 2026 as part of a seven-chip platformNVIDIA technical blog
March 2026LPX "will be available in the second half of 2026, coincident with the broader Vera Rubin rollout"; initial focus on model builders and service providers, not broad OEM availabilityStorageReview
Later 2026Vera Rubin itself is "ramping into full production with racks running at partners" (August 26, 2026). NVIDIA's release does not separately date LPX.NVIDIA Form 8-K

Honest summary: NVIDIA says it is for the second half of 2026 and for a narrow set of customers. NVIDIA's technical post and product page give no general availability date or price. Aquanode does not rent LPX or any LPU.

When this matters to you

  • You serve a very large model with strict per-user latency targets and run your own racks: LPX is built for you, and you would buy it through NVIDIA's channel.
  • You rent GPUs by the hour: nothing about LPX changes what you can deploy this quarter. Latency work on B300 or H200 (batching, speculative decoding, quantization, KV-cache management) is where gains are available now. See the B300 guide.
  • You track NVIDIA's roadmap: LPX signals that NVIDIA expects inference to be served by heterogeneous hardware, with HBM-based GPUs for capacity and SRAM-based chips for speed.

Cost

There is no public LPX price, and we do not type an hourly price for any GPU in this post. To estimate cost on hardware you can rent, take a cited tokens-per-second figure for your model, convert it to tokens per GPU-hour, and multiply by the live hourly price in the box below.

What to run today

The Groq 3 LPX is not something Aquanode rents, and this post is not a way to reserve it. Aquanode manages and optimizes GPUs for training and inference workloads, and you can rent the GPUs in the box below on demand to build and tune latency-sensitive serving now.

What's next

NVIDIA's pages we reviewed do not publish a roadmap beyond the second-half-2026 rollout for LPX. For the wider generation picture, see Rubin vs Blackwell vs Hopper.

FAQ

What is the NVIDIA Groq 3 LPU?

A language processing unit based on Groq's technology that NVIDIA licensed in December 2025. Each chip has 500 MB of on-chip SRAM and is used in 256-chip LPX racks for low-latency decode.

Did NVIDIA buy Groq?

No. Groq announced a non-exclusive licensing agreement on December 24, 2025. Key Groq leaders joined NVIDIA, and Groq remains a separate company. The reported $20 billion value comes from the press, not from the companies.

Is Groq 3 LPX a replacement for GPUs?

No. NVIDIA positions it as a companion to Vera Rubin NVL72, with Rubin GPUs and LPUs jointly computing each token.

Is Groq 3 LPX available?

NVIDIA's GTC 2026 announcement said the second half of 2026; StorageReview reports an initial focus on model builders and service providers. There is no public self-serve availability, and Aquanode does not rent it.

What is the difference between LPX and Rubin CPX?

CPX was announced in 2025 as a prefill (context) accelerator with GDDR7 and is not shipping. LPX is a decode accelerator built on SRAM.

Sources

#datacenter gpu#ai accelerators#nvidia rubin#groq lpu#nvidia lpu#low-latency inference

Submit the job. Everything after that is ours.

Sign up in 60 seconds. Pay for the GPU minutes you actually use.

© 2026 Aquanode. All rights reserved.

All trademarks, logos and brand names are the property of their respective owners.