Etched Sohu is an application-specific chip (ASIC) built to run transformer models only, announced in June 2024 with the claim that one 8-chip server matches 160 NVIDIA H100 GPUs. As of October 2026, Etched says it has shipped its first rack to a customer, but we found no independent benchmark of any Etched hardware, so every performance number below is Etched's claim.
TL;DR
- What it is: a chip that hard-wires the transformer architecture and drops general-purpose features, trading flexibility for inference throughput.
- What has shipped: Etched's own progress page says it shipped its first rack to Jane Street on August 18, 2026. That page does not use the name Sohu in the entry we read, and gives no rack specifications.
- What has not happened: no third-party benchmark, no MLPerf submission we could find, and no public price.
- The big numbers (20x an H100 server, 500,000 tokens per second on Llama 70B) date from 2024 marketing and have not been independently verified.
- Verdict: interesting bet, not something you can plan a production workload around today. Use NVIDIA GPUs for anything you need to run now.
What Etched claimed in 2024
When Etched announced Sohu in June 2024, Tom's Hardware reported the company's claims that a single 8xSohu server equals about 160 H100 GPUs and that H100s use only about 3.3% of their transistors on matrix multiplication, the core transformer operation. Electronics Weekly's June 2024 coverage quotes Etched as saying "over 500,000 tokens per second running Llama 70B" and that an "8xSohu server replaces 160 H100s", and says the chip is made on TSMC 4nm. A later review (fast.io, last reviewed June 29, 2026) says the figure "comes entirely from Etched's own marketing materials and demo videos." We found no memory capacity or bandwidth for Sohu on a page from Etched or major press, so we do not state one.
We could not retrieve Etched's original 2024 announcement page (the URL we tried returned not found), so we cannot quote the test conditions: model precision, batch size, context length or software baseline. Without them, the comparison cannot be reproduced.
Shipping status, with dates
This is the timeline we can source. Treat the first two as Etched statements and the last as an independent reading.
- 2024 to early 2025: initial shipments were targeted for this window, according to the fast.io review. It did not happen.
- Mid 2026: Etched's homepage says its A0 silicon came back from TSMC N4P, that it is validating its first rack-scale product with customers "to fulfill $1B in demand", and that its first racks ship "this summer". It also claims the system can run trillion-parameter sparse mixture-of-experts models at "80%+ Peak FLOPs without thermal throttling" and says more performance detail will come later.
- July 23, 2026: Etched's progress page lists a $300M raise at a $10.3B valuation. (The homepage banner and body quote other fundraising totals, so we do not repeat any funding figure as settled.)
- August 18, 2026: Etched's progress page says it shipped its first rack to Jane Street, "the first step on our mission to run the world's inference."
- June 2026 (independent): the fast.io review stated that, as of its June 2026 review, Sohu had not shipped to external customers and no independent benchmarks existed. The August entry post-dates it.
What we did not find: any customer, other than the named first rack, confirming it runs production traffic; any published tokens per second on a named model from Etched's current rack; any price.
One more caution. The 2026 pages describe a rack-scale product and refer to "frontier inference clusters". They do not tie the shipped rack to the 2024 Sohu spec sheet. Do not assume the 2024 performance claims describe what was delivered.
Spec table
| Spec | Etched Sohu (2024 claims) | NVIDIA H100 SXM | NVIDIA DGX B200 (8 GPUs) |
|---|---|---|---|
| Memory | not published | 80GB | 1,440 GB total |
| Memory bandwidth | not published | 3.35 TB/s | 64 TB/s total |
| Process | TSMC 4nm (Electronics Weekly, 2024); N4P per Etched's 2026 homepage | not covered here | not covered here |
| Peak compute | not published | 3,958 teraFLOPS FP8 (with sparsity) | 72 PFLOPS FP8, 144 PFLOPS FP4 (sparse) |
| Power | not published | up to 700W | about 14.3 kW max |
| Supported workloads | transformers only (Etched) | any | any |
Where the table says "not published", we did not find the number on a vendor page. The NVIDIA figures are from NVIDIA's own pages; the H100 compute number is marked "with sparsity" there.
Architecture: why a transformer-only chip could be fast
A GPU keeps flexible hardware for graphics, convolutions, recurrent nets and arbitrary kernels. Etched's argument is that nearly every frontier language model is a transformer, so a chip that spends its area on the matrix multiplies and attention those models need can use more of its silicon on useful work. That is the logic behind the 3.3% figure above.
The trade is risk. If the dominant architecture changes (state space models, diffusion language models, new attention variants), a chip fixed to today's transformer loses most of its advantage, while a GPU runs the new model next month. Etched also needs its own compiler and kernels; the ecosystem around Tensor Cores, CUDA and open-source servers is what NVIDIA customers rely on today.
Infrastructure needs
Etched has said its first product is a rack-scale system. We found no published power, cooling or networking figures. For comparison, NVIDIA's DGX B200 is listed at about 14.3 kW maximum, with 14.4 TB/s of aggregate NVLink bandwidth, and needs a matching network and cooling plan. See our guide to what NVLink is for why the interconnect matters at rack scale.
When to choose Etched versus NVIDIA
- Right now, for anything in production: NVIDIA. Etched has one named customer for its first rack and no independent data.
- If you are a very large inference operator evaluating hardware two to three years out: talk to Etched, ask for a reproducible benchmark with model, precision and batch size, and insist on a pilot.
- If you want low latency on supported models without owning hardware: compare with Cerebras, which sells a cloud inference service today.
- For picking among shipping NVIDIA parts, read H200 vs B200 vs GB200. The wider landscape is in the datacenter GPU overview.
Cost
There is no published Etched price or rental rate, so we do not state a cost per token. For the GPU side, take the tokens per second you measure on your own model and settings, multiply by 3,600 for tokens per GPU-hour, and compare with the live hourly price below. Be wary of any cost comparison that uses a 20x claim without stating the baseline.
What to run today
Etched hardware is not something you can rent on demand today. Aquanode manages and optimizes GPUs for training and inference workloads; you can rent the NVIDIA GPUs below on demand and benchmark your own model.
What's next
Etched's own pages say more performance details will come. When there is a reproducible benchmark with a named model and configuration, or an MLPerf Inference submission, the comparison becomes testable. On the NVIDIA side, see the B200 guide and the Rubin guide for what Etched will actually be compared against.
FAQ
Has Etched Sohu shipped?
Etched says it shipped its first rack to Jane Street on August 18, 2026. We found no independent confirmation of volume shipments, and a June 2026 review said it had not yet shipped to external customers.
Is Sohu really 20x faster than an H100?
That is Etched's 2024 claim (one 8xSohu server versus 160 H100s). It has not been independently verified, and the test conditions were not available to us.
Can Sohu run any model?
No. It is built for transformers only, per Etched. Other architectures are not supported by design.
Is there an Etched MLPerf result?
We did not find one.
Can I rent Sohu?
Not through Aquanode, and we found no public rental offering. The box above shows the GPUs you can rent today.
Sources
- Etched homepage (ship timing, N4P silicon, demand, performance claims): https://www.etched.com/
- Etched progress page (August 18, 2026 first rack shipment): https://www.etched.com/progress
- Tom's Hardware, Sohu announcement coverage (June 26, 2024): https://www.tomshardware.com/tech-industry/artificial-intelligence/sohu-ai-chip-claimed-to-run-models-20x-faster-and-cheaper-than-nvidia-h100-gpus
- Electronics Weekly, Etched chip an order of magnitude faster than Blackwell (June 2024; TSMC 4nm, 500,000 tokens per second, 160 H100s): https://www.electronicsweekly.com/news/business/etched-chip-an-order-of-magnitude-faster-than-blackwell-2024-06/
- Etched homepage (A0 silicon from TSMC N4P, first rack-scale product): https://www.etched.com/
- fast.io Etched review (last reviewed June 29, 2026; secondary source for shipping status and the "Etched's own marketing" characterization): https://www.fast.io/resources/etched-ai-review-2026.md
- NVIDIA H100: https://www.nvidia.com/en-us/data-center/h100/
- NVIDIA DGX B200: https://www.nvidia.com/en-us/data-center/dgx-b200/