Datacenter GPUs in 2026: Every AI Chip Compared

Back
Team Aquanode

Team Aquanode

Sarthak Vaish

Updated OCTOBER 8, 2026Published OCTOBER 8, 2026

The datacenter GPU market in October 2026 has four layers: NVIDIA Hopper and Blackwell, which you can rent today; NVIDIA Rubin and AMD's MI400 series, which are ramping; AMD's MI300 and MI350 families, which are shipping now; and a set of non-GPU accelerators from Google, AWS, Intel, Cerebras, Etched and Huawei. This page is the map. Every chip below links to a guide with the detail, and every number is tied to a vendor source listed at the end.

How to read it: memory and bandwidth decide which models fit and how fast they decode, the number format decides how much compute you get at a given precision, and the interconnect decides how far a job scales. Status tells you whether you can use the chip at all.

TL;DR

  • For most teams renting GPUs this year, the choice is among H100, H200, B200 and B300, plus AMD's MI300X and MI355X.
  • Memory has grown faster than anything else: from 80 GB on H100 to 141 GB on H200, 180 GB per GPU on B200, and 288 GB on Rubin. Bandwidth went from 3.35 TB/s on H100 to as much as 22 TB/s on Rubin, per NVIDIA's technical blog (its Vera Rubin NVL72 page lists 19.2 TB/s).
  • Rubin and AMD's MI455X are the next generation. NVIDIA says Rubin is in full production with partner products in the second half of 2026; AMD's MI455X was unveiled in July 2026.
  • Google TPU, AWS Trainium, Intel Gaudi 3, Cerebras, Etched and Huawei Ascend are real alternatives, but each is tied to its maker's cloud, software stack or region.
  • Rule of thumb: pick the smallest, oldest chip whose memory fits your model, then check the live price.

Every datacenter AI chip in one table

Status means: shipping (generally sold or deployed), ramping (the vendor says production has started, availability is limited), announced (specs published, not shipping to buyers). Figures are vendor-published; cells that say "not published" mean the vendor has not released the number.

ChipFamilyStatusMemoryHeadline specGuide
H100 SXMNVIDIA HopperShipping80 GB HBM33.35 TB/s bandwidthSXM vs NVL vs PCIe, /gpu/nvidia-h100
H100 NVLNVIDIA HopperShipping94 GBDual-card NVLink formSXM vs NVL vs PCIe, /gpu/nvidia-h100-nvl
H200NVIDIA HopperShipping141 GB HBM3e4.8 TB/s bandwidthH200 guide, SXM vs NVL vs PCIe, /gpu/nvidia-h200
GH200NVIDIA Grace HopperShippingNVL2: 288 GB HBMNVLink-C2C at 900 GB/sGH200 guide, /gpu/nvidia-gh200
H20NVIDIA Hopper (China)Shipping, region-limitedNot published by NVIDIAHopper-based variantH20 guide
B200NVIDIA BlackwellShipping180 GB HBM3e per GPU (1,440 GB across 8)8 TB/s per GPU (64 TB/s across 8)B200 guide, /gpu/nvidia-b200
B300NVIDIA Blackwell UltraShipping2.1 TB across 8 GPUs (DGX B300)108 PFLOPS dense FP4 per 8-GPU systemB300 guide, /gpu/nvidia-b300
GB200 NVL72NVIDIA Blackwell rackShipping13.4 TB HBM3E per rack576 TB/s GPU bandwidth, 720 PFLOPS dense FP4GB200 guide, /gpu/nvidia-gb200
GB300 NVL72NVIDIA Blackwell Ultra rackShipping20 TB per rackUp to 576 TB/s GPU bandwidthGB300 vs GB200, /gpu/nvidia-gb300
DGX SparkNVIDIA GB10 desktopShipping128 GB unified273 GB/s, 1 PFLOP FP4 claimDGX Spark guide
DGX StationNVIDIA GB300 desk-sideShipping via partners748 GB coherent (252 GB HBM3e)7.1 TB/s GPU bandwidthDGX Station guide
RubinNVIDIA RubinRamping (NVIDIA: full production)Up to 288 GB HBM419.2 to 22 TB/s (NVIDIA pages differ), 50 PFLOPS NVFP4Rubin guide, /gpu/nvidia-vera-rubin
Vera Rubin NVL72NVIDIA Rubin rackRamping72 Rubin GPUs216 to 260 TB/s rack NVLink (NVIDIA pages differ)Vera Rubin NVL72 guide
Rubin CPXNVIDIA RubinAnnounced128 GB GDDR730 PFLOPS NVFP4Rubin CPX guide
MI300XAMD CDNA 3Shipping192 GB HBM35.3 TB/s bandwidth/gpu/amd-mi300x-price, MI355X guide
MI325XAMD CDNA 3Shipping256 GB HBM3E6 TB/s bandwidthMI325X guide, /gpu/amd-mi325x
MI355XAMD CDNA 4ShippingUp to 288 GB HBM3E8 TB/s, FP4 and FP6 supportMI355X guide, /gpu/amd-mi355x
MI455X (MI400 series)AMD CDNA 5Launched July 2026, volume 2H 2026 (AMD)432 GB HBM440 PFLOPS FP4 (AMD's peak)MI400 and MI450 guide
TPU Ironwood (v7)GoogleAvailable on Google Cloud192 GB HBM7.37 TB/s, 4,614 TFLOPs FP8TPU vs GPU
Trainium3AWSAvailable on AWS144 GB HBM3e4.9 TB/sTrainium vs NVIDIA
Gaudi 3IntelShipping128 GB HBM2e3.7 TB/sGaudi 3 vs NVIDIA
WSE-3CerebrasShipping as systems44 GB on-chip SRAM21 PB/s on-chip bandwidthCerebras vs NVIDIA
SohuEtchedShipping status not confirmed by Etched144 GB HBM3E (reported)Transformer-only ASICEtched Sohu vs NVIDIA
Ascend 950PR / 950DTHuawei950PR launched; 950DT announced for Q4 2026128 GB / 144 GB in-house HBM1 PFLOPS FP8, 2 PFLOPS FP4Ascend vs NVIDIA

Sources for each row are in the family sections below and in the list at the end. The B200 per-GPU numbers are computed by dividing NVIDIA's DGX B200 totals (1,440 GB and 64 TB/s) by eight GPUs; B300 is shown as NVIDIA's system total because the per-GPU figure is not stated on the DGX page. FLOPS across rows use different precisions and sparsity settings, which is why the table leads with memory. See TFLOPS and FP4 for how to read the numbers.

NVIDIA Hopper: the workhorse generation

Hopper is still the volume generation. NVIDIA's H100 page lists 80 GB at 3.35 TB/s for the SXM part and 94 GB for the NVL variant. The H200 page lists 141 GB of HBM3e at 4.8 TB/s, which NVIDIA describes as nearly double H100's capacity and 1.4 times its bandwidth.

Choose Hopper when the model fits and price per hour matters more than speed.

NVIDIA Blackwell and Blackwell Ultra

Blackwell is the current generation for new deployments. DGX B200 has eight GPUs with 1,440 GB of HBM3e and 64 TB/s of memory bandwidth, and 144 PFLOPS sparse FP4 (72 dense), joined by fifth-generation NVLink. DGX B300 uses Blackwell Ultra with 2.1 TB of GPU memory and 144 PFLOPS sparse (108 dense) FP4.

At rack scale, NVIDIA's GB200 NVL72 page lists 72 Blackwell GPUs, 13.4 TB of HBM3E, 576 TB/s of GPU memory bandwidth and 720 PFLOPS dense FP4. The GB300 NVL72 page lists 72 Blackwell Ultra GPUs, 20 TB of GPU memory and up to 576 TB/s.

Two desk-side systems use the same architecture: the DGX Spark and the DGX Station GB300. Both are covered with rent-vs-buy math.

NVIDIA Rubin: announced, not the default

NVIDIA's January 5, 2026 announcement states: "NVIDIA Rubin is in full production, and Rubin-based products will be available from partners the second half of 2026" (NVIDIA newsroom). Per NVIDIA's technical blog, the Rubin GPU has up to 288 GB of HBM4, up to 22 TB/s of bandwidth (about 2.8 times Blackwell), and up to 50 PFLOPS of NVFP4. Vera Rubin NVL72 combines 72 Rubin GPUs and 36 Vera CPUs, with 216 to 260 TB/s of rack NVLink (NVIDIA's product page and technical blog differ).

"In production" is not the same as "rentable by the hour." Treat Rubin as a planning target until instances show live prices below.

AMD Instinct

AMD's line has moved quickly. MI300X carries 192 GB of HBM3 at 5.3 TB/s, MI325X 256 GB of HBM3E at 6 TB/s, and MI355X up to 288 GB of HBM3E at 8 TB/s with FP4 and FP6 support (Tom's Hardware, Next Platform).

  • AMD MI355X guide
  • AMD MI325X guide
  • AMD MI400 and MI450 guide: AMD unveiled the MI455X and Helios rack in July 2026, with 432 GB of HBM4 and 40 PFLOPS of FP4 per GPU as AMD's peak figures, and shipments reported for the end of Q3 2026 ramping through 2027 (StorageReview). We could not verify this against AMD's own page, so treat the timing as reported.

AMD's large memory pools suit models that barely fit on an 80 GB card. Software maturity is the main question; the guides cover ROCm support.

Interconnect and memory

A chip is only as useful as its links. These posts cover the parts that decide scaling:

Other accelerators

These chips are built by cloud and silicon companies rather than sold as general GPUs.

The common thread: each is excellent inside its maker's environment and harder to move workloads onto than a CUDA GPU.

The NVIDIA roadmap, dated only from sources

ItemWhat the vendor saidSource
RubinIn full production; partner products in the second half of 2026NVIDIA, January 5, 2026
Rubin CPXExpected at the end of 2026 (September 2025 announcement)NVIDIA
AMD MI455X and HeliosUnveiled July 2026; shipments reported from end of Q3 2026StorageReview
Huawei Ascend 950DTQ4 2026Huawei Connect 2025

How to choose

  1. Start from the model, not the chip. Weights at your precision plus KV cache must fit in memory. A 70B model at 8-bit needs about 70 GB for weights alone (70 billion x 1 byte), so it fits on an H100 with little room; an H200 or MI300X leaves room for context. Our KV cache and quantization entries explain the sizing.
  2. Decode speed follows bandwidth. Generating tokens reads the weights each step, so higher TB/s means faster output. That is the real gap between H100, H200 and B200.
  3. Training scale follows interconnect. If a job spans many GPUs, NVLink inside the node and the network outside it matter more than single-chip FLOPS.
  4. Mixture-of-experts changes the math. Memory capacity decides what loads; only active experts move per token. See mixture-of-experts.
  5. Pick the oldest chip that fits. If an H100 does the job, a B300 only adds cost. If it does not fit, step up one tier, not three.
  6. Watch the software. CUDA runs everywhere; ROCm, TPU, Trainium and Gaudi each need porting work.
  7. Desk or datacenter? For prototyping and privacy, a DGX Spark can be enough; for anything sustained or time-bound, rent. Compare chips directly on our compare pages and the GPU index.

Rent today

Aquanode manages and optimizes GPUs for training and inference workloads, and you can rent the GPUs in the box below on demand. Prices update live, and a row reads "None right now" when no offer is available.

For full price history and specs, see /pricing and the GPU index.

Related reading: who competes with NVIDIA in AI, MI300X vs H100 and H200 for inference, AMD vs NVIDIA GPUs for AI and the best GPUs for AI.

FAQ

What is the best datacenter GPU for AI in 2026?

There is no single best. For capacity and speed, Blackwell (B200, B300) leads what you can rent. For cost-sensitive work where the model fits, H100 and H200 remain widely used. The smallest chip whose memory fits your model is usually the right start.

How much memory does each current NVIDIA datacenter GPU have?

H100 SXM has 80 GB, H200 has 141 GB, a B200 works out to 180 GB per GPU (NVIDIA lists 1,440 GB for eight), and Rubin is specified at up to 288 GB of HBM4.

Is NVIDIA Rubin available?

NVIDIA says Rubin is in full production and partner products arrive in the second half of 2026 (January 5, 2026 announcement). Check the live box above for what can be rented today.

How does AMD compare with NVIDIA?

AMD offers more memory per chip at its tier: 192 GB on MI300X, 256 GB on MI325X, up to 288 GB on MI355X. NVIDIA has the broader software ecosystem. See the AMD guides above.

Are TPUs and Trainium alternatives to GPUs?

Yes, inside Google Cloud and AWS respectively. They are not general rentable GPUs, and workloads usually need porting.

What does HBM have to do with GPU choice?

HBM capacity decides which models fit, and its bandwidth sets decode speed. See HBM and HBM3e vs HBM4.

Sources

#datacenter gpu#nvidia blackwell#nvidia hopper#nvidia rubin#amd instinct#gpu interconnect#ai accelerators#gpu roadmap

Submit the job. Everything after that is ours.

Sign up in 60 seconds. Pay for the GPU minutes you actually use.

© 2026 Aquanode. All rights reserved.

All trademarks, logos and brand names are the property of their respective owners.