Tenstorrent builds AI accelerators and the computers around them, based on RISC-V cores and an open-source software stack. Its Blackhole cards are listed at $999 to $1,399 on its own site, with 120 Tensix cores and 28 to 32 GB of GDDR6, which makes them a buy-and-own developer platform rather than a drop-in replacement for an NVIDIA H100 or B200.
TL;DR
- What Tenstorrent does: it designs AI chips and sells them as PCIe cards, desktop-class QuietBox workstations and the 32-chip Galaxy server. Its CEO is Jim Keller, per the company's about page.
- Blackhole card specs from Tenstorrent: 120 Tensix cores, 16 "big" RISC-V cores, 180MB SRAM, 28 or 32 GB GDDR6, 664 TFLOPS in BLOCKFP8, 300W.
- It has no HBM and no NVLink. Memory bandwidth is 448 to 512 GB/s per card, against 3.35 TB/s on an H100.
- Its software is open source (Apache 2.0 for TT-Metalium and TT-NN), but Tenstorrent's own repo says Blackhole software optimization is still under active development.
- Verdict: interesting for developers, researchers and teams who want open hardware and low entry prices. For production LLM training or serving at scale, NVIDIA GPUs are the practical choice today.
What Tenstorrent builds
Tenstorrent's homepage says it "builds computers for AI". The company's about page describes open architectures and open-source software, and names a RISC-V CPU effort. Its product families are:
- Wormhole: the earlier generation (cards n150 and n300, listed at $999 to $1,449).
- Blackhole: the current card generation (p100a, p150a, p150b).
- QuietBox: workstations built from Tenstorrent cards.
- Galaxy: a 32-chip server aimed at rack deployments.
Each chip is built from Tensix cores, each with its own small CPU cores and local SRAM, connected by an on-chip network. Blackhole cards also carry Ethernet-based links, so cards connect to each other directly instead of through a separate switch ASIC.
Blackhole spec table
All Tenstorrent figures are from its Blackhole hardware page. NVIDIA figures are from NVIDIA's pages.
| Spec | Blackhole p100a | Blackhole p150a | NVIDIA H100 SXM | NVIDIA DGX B200 (8 GPUs) |
|---|---|---|---|---|
| List price | $999 | $1,399 | not listed | not listed |
| Compute units | 120 Tensix cores | 120 Tensix cores | not covered here | not covered here |
| Memory | 28 GB GDDR6 | 32 GB GDDR6 | 80GB HBM3 | 1,440 GB HBM3e total |
| Memory bandwidth | 448 GB/s | 512 GB/s | 3.35 TB/s | 64 TB/s total |
| On-chip SRAM | 180MB | 180MB | not covered here | not covered here |
| Compute | 664 TFLOPS BLOCKFP8 | 664 TFLOPS BLOCKFP8 | 3,958 teraFLOPS FP8 (with sparsity) | 72 PFLOPS FP8 (sparse) |
| Power | 300W | 300W | up to 700W | about 14.3 kW max |
| Card links | none | 4x QSFP-DD 800G | NVLink 900GB/s | 14.4 TB/s aggregate NVLink |
Do not compare the compute rows directly. BLOCKFP8 is a block floating point format, and NVIDIA's figure is marked "with sparsity" on its page. The memory rows are the fair comparison: a Blackhole card has about one seventh of an H100's bandwidth (512 GB/s against 3.35 TB/s, our arithmetic). Token generation for large language models is mostly limited by memory bandwidth, so that gap matters. See our HBM glossary entry for why.
Galaxy and QuietBox
Tenstorrent lists the Galaxy Blackhole system with:
- 32 Blackhole ASICs, 23 PFLOPS at Block FP8
- 6.2 GB of on-chip SRAM at 2.9 PB/s, and 1 TB of GDDR6 at 16 TB/s
- 8 to 10 kW average power, 12 kW max, configurable up to 14.5 kW
- 10x 400 GbE links per ASIC (32 TB/s total fabric), and up to 56x 800 GbE ports for scale-out
- $160,000 list price
For scale, the DGX B200 listed on NVIDIA's page has 1,440 GB of HBM3e at 64 TB/s total and about 14.3 kW maximum. Galaxy has less than 1,440 GB of memory at roughly a quarter of the bandwidth, at a similar power draw (our arithmetic from the two pages). Tenstorrent has not published an MLPerf result that we found, and we did not find independent throughput numbers, so we cannot say what that means per token.
QuietBox prices on Tenstorrent's page when we looked: $9,999 for the QuietBox 2 (Blackhole), $11,999 for the original Blackhole QuietBox (sold out), and $15,000 for the Wormhole QuietBox (sold out). The page was inconsistent about the QuietBox 2's chip count (it mentioned both four processors and two p300c cards) and gave no power figure, so confirm on Tenstorrent's site.
Software
Tenstorrent's open-source repo describes TT-Metalium as a low-level programming model for kernel development and TT-NN as a Python and C++ neural network operator library, both under the Apache 2.0 license. The repo supports vLLM through a Tenstorrent plugin, and its model tables list Wormhole systems and the p150 Blackhole. It also states that Blackhole software optimization is under active development.
That is the main trade. NVIDIA's CUDA stack means almost any model and framework runs the day it is released. On Tenstorrent you work with a smaller set of tuned models, and you can read and change the stack, which some teams value. The general idea behind Tensix-style compute is covered in our Tensor Core glossary entry, though the architecture differs.
Infrastructure needs
- Blackhole card: 300W, PCIe 5.0 x16, active or passive cooling by model. It fits a workstation or server.
- Galaxy: up to 14.5 kW, high-speed Ethernet, a rack with proper power and cooling.
- Compared with NVIDIA's liquid-cooled rack-scale systems such as the GB200 NVL72, Tenstorrent's footprint is much smaller, as is its ceiling.
When to choose which
- Learning, research, porting and kernel work on non-CUDA hardware: Tenstorrent cards are one of the few ways to own an open accelerator for about $1,000 to $1,400.
- Running a small model at home or in a lab with modest speed needs: possible, but expect to work with the supported model list.
- Training or serving large models: NVIDIA. HBM bandwidth, NVLink and the CUDA ecosystem are the reasons; see H200 vs B200 vs GB200 and the datacenter GPU overview.
- Looking for other non-NVIDIA inference silicon: compare Cerebras, Etched and TPU.
Cost
Tenstorrent publishes purchase prices (above), not rental rates, and we found no sourced throughput per card, so we do not convert them to tokens per dollar. For GPUs on Aquanode, measure tokens per second on your own model, multiply by 3,600 for tokens per GPU-hour, and compare it with the live hourly price below.
Rent today
Tenstorrent hardware is bought, not rented, and it is not in the box. Aquanode manages and optimizes GPUs for training and inference workloads; you can rent the NVIDIA GPUs below on demand.
What's next
Tenstorrent's repo says Blackhole optimization is in progress, so expect performance to move with software. For the NVIDIA parts a buyer would otherwise consider, read the B200 guide.
FAQ
What does Tenstorrent do?
It designs AI accelerators and sells them as cards, workstations and servers, with RISC-V cores and open-source software, led by CEO Jim Keller according to its about page.
How much does a Tenstorrent Blackhole cost?
Tenstorrent lists the p100a at $999 and the p150a and p150b at $1,399, as of our visit to its page in October 2026.
Does Blackhole use HBM?
No. The cards use GDDR6, 28 or 32 GB at 448 to 512 GB/s.
Is Tenstorrent faster than an H100?
We found no independent benchmark. On paper, the H100 has about seven times the memory bandwidth of a p150a, and the compute figures use different formats.
Is the software open source?
TT-Metalium and TT-NN are Apache 2.0 licensed per Tenstorrent's repository.
Sources
- Tenstorrent Blackhole cards: https://tenstorrent.com/en/hardware/blackhole
- Tenstorrent Galaxy: https://tenstorrent.com/en/hardware/galaxy
- Tenstorrent QuietBox: https://tenstorrent.com/en/hardware/tt-quietbox
- Tenstorrent about page: https://tenstorrent.com/en/about
- Tenstorrent tt-metal repository: https://github.com/tenstorrent/tt-metal
- NVIDIA H100: https://www.nvidia.com/en-us/data-center/h100/
- NVIDIA DGX B200: https://www.nvidia.com/en-us/data-center/dgx-b200/