Cerebras vs NVIDIA: WSE-3 vs H100 and B200 (2026)

Back
Team Aquanode

Team Aquanode

Sarthak Vaish

Updated OCTOBER 8, 2026Published OCTOBER 8, 2026

Cerebras builds one processor out of an entire silicon wafer, while NVIDIA builds many small GPUs and links them with NVLink and networking. The wafer gives Cerebras very fast on-chip memory and very high single-user token speed on models that fit its system, but you cannot rent it like a GPU, and most of its speed numbers are Cerebras's own.

TL;DR

  • The Cerebras WSE-3 is one chip with 4 trillion transistors, 900,000 cores and 44GB of on-chip SRAM. An NVIDIA H100 has 80GB of HBM and a B200 node has far more total memory, but at much lower bandwidth per byte of capacity.
  • Cerebras sells fast tokens per user. NVIDIA sells a general platform that trains, fine-tunes and serves almost any model, with the largest software ecosystem.
  • Cerebras's headline speedups (16x, 30x) are vendor claims, mostly measured against older or unnamed GPU setups. Treat them as a direction, not a verdict on B200.
  • Groq's LPU is the other SRAM-first inference design. NVIDIA signed a licensing deal with Groq in December 2025.
  • Verdict: choose Cerebras for very low latency on a supported open model via its cloud. Choose NVIDIA GPUs for training, fine-tuning, custom models and anything you need to run yourself.

Spec comparison

SpecCerebras WSE-3Cerebras WSE-3T (CS-4)NVIDIA H100 SXMNVIDIA DGX B200 (8 GPUs)
Process5nm (TSMC)not published in the releasenot covered herenot covered here
Transistors4 trillion4 trillionnot covered herenot covered here
Cores900,000900,000not covered herenot covered here
Memory44GB on-chip SRAM44GB on-chip SRAM80GB HBM31,440 GB HBM3e total
Memory bandwidth21 PB/s43.2 PB/s3.35 TB/s64 TB/s total
Peak AI compute125 PFLOPS250 PFLOPS3,958 TFLOPS FP8 (with sparsity)72 PFLOPS FP8 and 144 PFLOPS FP4 (sparse)
Powernot published on the CS-3 pagenot published in the releaseup to 700Wabout 14.3 kW max

Read the table with care. Cerebras's peak compute is its own headline figure and does not state precision or sparsity in the sources we used. NVIDIA's figures are marked "with sparsity" on its pages, so the two columns are not apples to apples. Memory capacity is the sharper difference: 44GB of SRAM holds only a small model, so Cerebras serves bigger models by streaming weights from external memory or by linking several wafers.

Architecture

Wafer-scale versus GPU

A normal chip is cut from a wafer into many dies. Cerebras keeps the whole wafer as one die of 46,225 mm2, so cores talk to each other over on-wafer wiring instead of leaving the package. Cerebras says this is why its on-chip bandwidth is 21 PB/s on the CS-3 page, which it describes as 2,625x more than GPU-based systems (a vendor claim). An H100 reads weights from HBM at 3.35 TB/s, and that bandwidth is what limits token speed at small batch sizes.

Systems

The CS-3 is the system built around one WSE-3. Cerebras's March 2024 launch coverage said clusters of up to 2,048 systems are possible, and that a 2,048-chip cluster can train Llama 2 70B in under a day (vendor claim).

In August 2026 Cerebras announced the CS-4, which uses three WSE-3T wafers per system and lists 750 PFLOPS of AI compute and 129.6 PB/s of memory bandwidth. Cerebras says first shipments begin in the quarter of the announcement. That is a vendor statement; we have not seen independent confirmation of shipments.

Performance: what is published

Every number here is from Cerebras unless stated.

  • October 2024: Cerebras reported 2,100 tokens per second on Llama 3.2 70B, which it called 16x faster than any known GPU solution and 68x faster than hyperscale clouds, citing Artificial Analysis. The comparison was against Hopper-era setups, not Blackwell.
  • August 2026, CS-4: Cerebras says it delivers more than 4,400 tokens per second per user on GPT-OSS-120B, up to 30x more than GPU solutions in a head-to-head with identical prompts. The release does not name the GPUs, configuration or precision, and it says results vary by model, context length and serving setup. It also says the CS-4 is up to 2x faster than the CS-3 and gives up to 10x more throughput per watt (vendor claims).

What is missing: a published MLPerf result for Cerebras systems in the sources we reviewed, and any independent head-to-head against a B200 system with a stated configuration. Until one exists, "30x" tells you that Cerebras is very fast per user, not how it compares on cost per million tokens against a tuned B200 deployment.

For the GPU side, NVIDIA's own pages are the source for per-chip specs. See our B200 guide and the H100 page for details, and the B200 vs H100 comparison for a side-by-side.

Where Groq fits

Groq's LPU (language processing unit) is also an SRAM-first inference chip. Cerebras's own comparison page says each Groq chip has 230MB of SRAM, so large models need many linked chips, and claims it is more than 6x faster on gpt-oss-120B (about 3,000 versus about 493 tokens per second, citing Artificial Analysis). Those are Cerebras's claims about a competitor, published on its own site.

In December 2025, Groq said it entered a non-exclusive licensing agreement with NVIDIA, and its founder Jonathan Ross and president Sunny Madra moved to NVIDIA while Groq continued as an independent company. A $20 billion price was reported by CNBC from a source and was not disclosed by the companies. So the "Groq vs NVIDIA" question is now partly "NVIDIA plus Groq technology", and we would not read old Groq-versus-GPU comparisons as current.

Business status

Cerebras filed to go public in 2026. Secondary coverage reports a Nasdaq listing under CBRS priced at $185 per share, raising about $5.55 billion, and a reported OpenAI agreement for up to 750 megawatts of compute through 2028. The same coverage says G42-related customers made up roughly 86% of 2025 revenue. We have not read the official pricing release, so treat the exact terms as reported, not confirmed. The point for engineers is customer concentration: the platform's roadmap follows a few very large buyers.

Infrastructure needs

  • Cerebras: you normally do not run it yourself. Cerebras sells access through its cloud and also sells systems; the software stack is Cerebras's own compiler and serving layer, with a limited list of supported models.
  • NVIDIA: a DGX B200 draws about 14.3 kW at maximum and has 14.4 TB/s of aggregate NVLink bandwidth per NVIDIA's page, so it needs liquid or high-capacity air cooling and a fast network. The H100 is up to 700W per GPU.
  • Software: CUDA, PyTorch, vLLM and SGLang run on NVIDIA out of the box. Anything custom (new architecture, LoRA stack, custom kernels) is a GPU job.

When to choose which

  • Latency-critical chat or agents on a supported open model, where tokens per second per user is the product: evaluate Cerebras inference and measure with your prompts.
  • Training, fine-tuning or running your own checkpoints: NVIDIA GPUs. See H200 vs B200 vs GB200 for picking a generation.
  • Batch throughput at the lowest cost per token: this depends on your model and batch size. Measure on a B200 or H100 before assuming a wafer wins; tokens per second per user and tokens per second per dollar are different goals.
  • Quantization matters on both sides. Our FP8 and FP4 glossary entries explain what the precision labels on the benchmarks mean.

Cost

We do not publish a Cerebras cost per token here because we have no sourced price. For GPUs, compute it yourself: take the tokens per second you measure on your model, multiply by 3,600 for tokens per GPU-hour, then divide the live hourly price below by that number. Measure with your own prompt lengths; vendor tokens per second figures use their own settings.

Rent today

Aquanode manages and optimizes GPUs for training and inference workloads. You can rent the NVIDIA GPUs below on demand, and the box shows what is available right now.

What's next

Cerebras's CS-4 is the successor to the CS-3 per its August 2026 announcement. On NVIDIA's side, the Rubin generation is the next step; see our Rubin guide and the datacenter GPU overview. For other non-NVIDIA options, read TPU vs GPU and Etched Sohu vs NVIDIA.

FAQ

Is Cerebras faster than NVIDIA?

Per user on supported models, Cerebras says yes, by large margins, and its own and Artificial Analysis-cited figures support high single-stream speed. The published comparisons mostly use older or unnamed GPU setups, so it is not established against a tuned B200 system.

Can I rent a Cerebras chip like a GPU?

Not as a GPU. Cerebras sells inference through its cloud and sells CS systems. It is not in the Aquanode box above.

Is Groq still independent?

Groq said it continues to operate independently after a non-exclusive licensing agreement with NVIDIA in December 2025, with its founder and president joining NVIDIA.

Does 44GB of SRAM limit model size?

On one wafer, yes. Cerebras says larger models are handled with external memory and multiple systems, and its CS-3 page claims scaling to 24 trillion parameter models on a single logical device. That is a vendor claim.

Should I train on Cerebras or on GPUs?

If you rely on open tooling, custom code and existing checkpoints, GPUs are the default. Cerebras makes sense when it supports your exact model and you value its speed.

Sources

#datacenter gpu#ai accelerators#cerebras#wafer-scale#nvidia h100#nvidia b200

Submit the job. Everything after that is ours.

Sign up in 60 seconds. Pay for the GPU minutes you actually use.

© 2026 Aquanode. All rights reserved.

All trademarks, logos and brand names are the property of their respective owners.