Tenstorrent Blackhole vs NVIDIA: Specs, Prices (2026)

Back
Team Aquanode

Team Aquanode

Sarthak Vaish

Updated OCTOBER 8, 2026Published OCTOBER 8, 2026

Tenstorrent builds AI accelerators and the computers around them, based on RISC-V cores and an open-source software stack. Its Blackhole cards are listed at $999 to $1,399 on its own site, with 120 Tensix cores and 28 to 32 GB of GDDR6, which makes them a buy-and-own developer platform rather than a drop-in replacement for an NVIDIA H100 or B200.

TL;DR

  • What Tenstorrent does: it designs AI chips and sells them as PCIe cards, desktop-class QuietBox workstations and the 32-chip Galaxy server. Its CEO is Jim Keller, per the company's about page.
  • Blackhole card specs from Tenstorrent: 120 Tensix cores, 16 "big" RISC-V cores, 180MB SRAM, 28 or 32 GB GDDR6, 664 TFLOPS in BLOCKFP8, 300W.
  • It has no HBM and no NVLink. Memory bandwidth is 448 to 512 GB/s per card, against 3.35 TB/s on an H100.
  • Its software is open source (Apache 2.0 for TT-Metalium and TT-NN), but Tenstorrent's own repo says Blackhole software optimization is still under active development.
  • Verdict: interesting for developers, researchers and teams who want open hardware and low entry prices. For production LLM training or serving at scale, NVIDIA GPUs are the practical choice today.

What Tenstorrent builds

Tenstorrent's homepage says it "builds computers for AI". The company's about page describes open architectures and open-source software, and names a RISC-V CPU effort. Its product families are:

  • Wormhole: the earlier generation (cards n150 and n300, listed at $999 to $1,449).
  • Blackhole: the current card generation (p100a, p150a, p150b).
  • QuietBox: workstations built from Tenstorrent cards.
  • Galaxy: a 32-chip server aimed at rack deployments.

Each chip is built from Tensix cores, each with its own small CPU cores and local SRAM, connected by an on-chip network. Blackhole cards also carry Ethernet-based links, so cards connect to each other directly instead of through a separate switch ASIC.

Blackhole spec table

All Tenstorrent figures are from its Blackhole hardware page. NVIDIA figures are from NVIDIA's pages.

SpecBlackhole p100aBlackhole p150aNVIDIA H100 SXMNVIDIA DGX B200 (8 GPUs)
List price$999$1,399not listednot listed
Compute units120 Tensix cores120 Tensix coresnot covered herenot covered here
Memory28 GB GDDR632 GB GDDR680GB HBM31,440 GB HBM3e total
Memory bandwidth448 GB/s512 GB/s3.35 TB/s64 TB/s total
On-chip SRAM180MB180MBnot covered herenot covered here
Compute664 TFLOPS BLOCKFP8664 TFLOPS BLOCKFP83,958 teraFLOPS FP8 (with sparsity)72 PFLOPS FP8 (sparse)
Power300W300Wup to 700Wabout 14.3 kW max
Card linksnone4x QSFP-DD 800GNVLink 900GB/s14.4 TB/s aggregate NVLink

Do not compare the compute rows directly. BLOCKFP8 is a block floating point format, and NVIDIA's figure is marked "with sparsity" on its page. The memory rows are the fair comparison: a Blackhole card has about one seventh of an H100's bandwidth (512 GB/s against 3.35 TB/s, our arithmetic). Token generation for large language models is mostly limited by memory bandwidth, so that gap matters. See our HBM glossary entry for why.

Galaxy and QuietBox

Tenstorrent lists the Galaxy Blackhole system with:

  • 32 Blackhole ASICs, 23 PFLOPS at Block FP8
  • 6.2 GB of on-chip SRAM at 2.9 PB/s, and 1 TB of GDDR6 at 16 TB/s
  • 8 to 10 kW average power, 12 kW max, configurable up to 14.5 kW
  • 10x 400 GbE links per ASIC (32 TB/s total fabric), and up to 56x 800 GbE ports for scale-out
  • $160,000 list price

For scale, the DGX B200 listed on NVIDIA's page has 1,440 GB of HBM3e at 64 TB/s total and about 14.3 kW maximum. Galaxy has less than 1,440 GB of memory at roughly a quarter of the bandwidth, at a similar power draw (our arithmetic from the two pages). Tenstorrent has not published an MLPerf result that we found, and we did not find independent throughput numbers, so we cannot say what that means per token.

QuietBox prices on Tenstorrent's page when we looked: $9,999 for the QuietBox 2 (Blackhole), $11,999 for the original Blackhole QuietBox (sold out), and $15,000 for the Wormhole QuietBox (sold out). The page was inconsistent about the QuietBox 2's chip count (it mentioned both four processors and two p300c cards) and gave no power figure, so confirm on Tenstorrent's site.

Software

Tenstorrent's open-source repo describes TT-Metalium as a low-level programming model for kernel development and TT-NN as a Python and C++ neural network operator library, both under the Apache 2.0 license. The repo supports vLLM through a Tenstorrent plugin, and its model tables list Wormhole systems and the p150 Blackhole. It also states that Blackhole software optimization is under active development.

That is the main trade. NVIDIA's CUDA stack means almost any model and framework runs the day it is released. On Tenstorrent you work with a smaller set of tuned models, and you can read and change the stack, which some teams value. The general idea behind Tensix-style compute is covered in our Tensor Core glossary entry, though the architecture differs.

Infrastructure needs

  • Blackhole card: 300W, PCIe 5.0 x16, active or passive cooling by model. It fits a workstation or server.
  • Galaxy: up to 14.5 kW, high-speed Ethernet, a rack with proper power and cooling.
  • Compared with NVIDIA's liquid-cooled rack-scale systems such as the GB200 NVL72, Tenstorrent's footprint is much smaller, as is its ceiling.

When to choose which

  • Learning, research, porting and kernel work on non-CUDA hardware: Tenstorrent cards are one of the few ways to own an open accelerator for about $1,000 to $1,400.
  • Running a small model at home or in a lab with modest speed needs: possible, but expect to work with the supported model list.
  • Training or serving large models: NVIDIA. HBM bandwidth, NVLink and the CUDA ecosystem are the reasons; see H200 vs B200 vs GB200 and the datacenter GPU overview.
  • Looking for other non-NVIDIA inference silicon: compare Cerebras, Etched and TPU.

Cost

Tenstorrent publishes purchase prices (above), not rental rates, and we found no sourced throughput per card, so we do not convert them to tokens per dollar. For GPUs on Aquanode, measure tokens per second on your own model, multiply by 3,600 for tokens per GPU-hour, and compare it with the live hourly price below.

Rent today

Tenstorrent hardware is bought, not rented, and it is not in the box. Aquanode manages and optimizes GPUs for training and inference workloads; you can rent the NVIDIA GPUs below on demand.

What's next

Tenstorrent's repo says Blackhole optimization is in progress, so expect performance to move with software. For the NVIDIA parts a buyer would otherwise consider, read the B200 guide.

FAQ

What does Tenstorrent do?

It designs AI accelerators and sells them as cards, workstations and servers, with RISC-V cores and open-source software, led by CEO Jim Keller according to its about page.

How much does a Tenstorrent Blackhole cost?

Tenstorrent lists the p100a at $999 and the p150a and p150b at $1,399, as of our visit to its page in October 2026.

Does Blackhole use HBM?

No. The cards use GDDR6, 28 or 32 GB at 448 to 512 GB/s.

Is Tenstorrent faster than an H100?

We found no independent benchmark. On paper, the H100 has about seven times the memory bandwidth of a p150a, and the compute figures use different formats.

Is the software open source?

TT-Metalium and TT-NN are Apache 2.0 licensed per Tenstorrent's repository.

Sources

#datacenter gpu#ai accelerators#tenstorrent#tenstorrent blackhole#risc-v#nvidia h100

Submit the job. Everything after that is ours.

Sign up in 60 seconds. Pay for the GPU minutes you actually use.

© 2026 Aquanode. All rights reserved.

All trademarks, logos and brand names are the property of their respective owners.