The datacenter GPU market in October 2026 has four layers: NVIDIA Hopper and Blackwell, which you can rent today; NVIDIA Rubin and AMD's MI400 series, which are ramping; AMD's MI300 and MI350 families, which are shipping now; and a set of non-GPU accelerators from Google, AWS, Intel, Cerebras, Etched and Huawei. This page is the map. Every chip below links to a guide with the detail, and every number is tied to a vendor source listed at the end.
How to read it: memory and bandwidth decide which models fit and how fast they decode, the number format decides how much compute you get at a given precision, and the interconnect decides how far a job scales. Status tells you whether you can use the chip at all.
TL;DR
- For most teams renting GPUs this year, the choice is among H100, H200, B200 and B300, plus AMD's MI300X and MI355X.
- Memory has grown faster than anything else: from 80 GB on H100 to 141 GB on H200, 180 GB per GPU on B200, and 288 GB on Rubin. Bandwidth went from 3.35 TB/s on H100 to as much as 22 TB/s on Rubin, per NVIDIA's technical blog (its Vera Rubin NVL72 page lists 19.2 TB/s).
- Rubin and AMD's MI455X are the next generation. NVIDIA says Rubin is in full production with partner products in the second half of 2026; AMD's MI455X was unveiled in July 2026.
- Google TPU, AWS Trainium, Intel Gaudi 3, Cerebras, Etched and Huawei Ascend are real alternatives, but each is tied to its maker's cloud, software stack or region.
- Rule of thumb: pick the smallest, oldest chip whose memory fits your model, then check the live price.
Every datacenter AI chip in one table
Status means: shipping (generally sold or deployed), ramping (the vendor says production has started, availability is limited), announced (specs published, not shipping to buyers). Figures are vendor-published; cells that say "not published" mean the vendor has not released the number.
| Chip | Family | Status | Memory | Headline spec | Guide |
|---|---|---|---|---|---|
| H100 SXM | NVIDIA Hopper | Shipping | 80 GB HBM3 | 3.35 TB/s bandwidth | SXM vs NVL vs PCIe, /gpu/nvidia-h100 |
| H100 NVL | NVIDIA Hopper | Shipping | 94 GB | Dual-card NVLink form | SXM vs NVL vs PCIe, /gpu/nvidia-h100-nvl |
| H200 | NVIDIA Hopper | Shipping | 141 GB HBM3e | 4.8 TB/s bandwidth | H200 guide, SXM vs NVL vs PCIe, /gpu/nvidia-h200 |
| GH200 | NVIDIA Grace Hopper | Shipping | NVL2: 288 GB HBM | NVLink-C2C at 900 GB/s | GH200 guide, /gpu/nvidia-gh200 |
| H20 | NVIDIA Hopper (China) | Shipping, region-limited | Not published by NVIDIA | Hopper-based variant | H20 guide |
| B200 | NVIDIA Blackwell | Shipping | 180 GB HBM3e per GPU (1,440 GB across 8) | 8 TB/s per GPU (64 TB/s across 8) | B200 guide, /gpu/nvidia-b200 |
| B300 | NVIDIA Blackwell Ultra | Shipping | 2.1 TB across 8 GPUs (DGX B300) | 108 PFLOPS dense FP4 per 8-GPU system | B300 guide, /gpu/nvidia-b300 |
| GB200 NVL72 | NVIDIA Blackwell rack | Shipping | 13.4 TB HBM3E per rack | 576 TB/s GPU bandwidth, 720 PFLOPS dense FP4 | GB200 guide, /gpu/nvidia-gb200 |
| GB300 NVL72 | NVIDIA Blackwell Ultra rack | Shipping | 20 TB per rack | Up to 576 TB/s GPU bandwidth | GB300 vs GB200, /gpu/nvidia-gb300 |
| DGX Spark | NVIDIA GB10 desktop | Shipping | 128 GB unified | 273 GB/s, 1 PFLOP FP4 claim | DGX Spark guide |
| DGX Station | NVIDIA GB300 desk-side | Shipping via partners | 748 GB coherent (252 GB HBM3e) | 7.1 TB/s GPU bandwidth | DGX Station guide |
| Rubin | NVIDIA Rubin | Ramping (NVIDIA: full production) | Up to 288 GB HBM4 | 19.2 to 22 TB/s (NVIDIA pages differ), 50 PFLOPS NVFP4 | Rubin guide, /gpu/nvidia-vera-rubin |
| Vera Rubin NVL72 | NVIDIA Rubin rack | Ramping | 72 Rubin GPUs | 216 to 260 TB/s rack NVLink (NVIDIA pages differ) | Vera Rubin NVL72 guide |
| Rubin CPX | NVIDIA Rubin | Announced | 128 GB GDDR7 | 30 PFLOPS NVFP4 | Rubin CPX guide |
| MI300X | AMD CDNA 3 | Shipping | 192 GB HBM3 | 5.3 TB/s bandwidth | /gpu/amd-mi300x-price, MI355X guide |
| MI325X | AMD CDNA 3 | Shipping | 256 GB HBM3E | 6 TB/s bandwidth | MI325X guide, /gpu/amd-mi325x |
| MI355X | AMD CDNA 4 | Shipping | Up to 288 GB HBM3E | 8 TB/s, FP4 and FP6 support | MI355X guide, /gpu/amd-mi355x |
| MI455X (MI400 series) | AMD CDNA 5 | Launched July 2026, volume 2H 2026 (AMD) | 432 GB HBM4 | 40 PFLOPS FP4 (AMD's peak) | MI400 and MI450 guide |
| TPU Ironwood (v7) | Available on Google Cloud | 192 GB HBM | 7.37 TB/s, 4,614 TFLOPs FP8 | TPU vs GPU | |
| Trainium3 | AWS | Available on AWS | 144 GB HBM3e | 4.9 TB/s | Trainium vs NVIDIA |
| Gaudi 3 | Intel | Shipping | 128 GB HBM2e | 3.7 TB/s | Gaudi 3 vs NVIDIA |
| WSE-3 | Cerebras | Shipping as systems | 44 GB on-chip SRAM | 21 PB/s on-chip bandwidth | Cerebras vs NVIDIA |
| Sohu | Etched | Shipping status not confirmed by Etched | 144 GB HBM3E (reported) | Transformer-only ASIC | Etched Sohu vs NVIDIA |
| Ascend 950PR / 950DT | Huawei | 950PR launched; 950DT announced for Q4 2026 | 128 GB / 144 GB in-house HBM | 1 PFLOPS FP8, 2 PFLOPS FP4 | Ascend vs NVIDIA |
Sources for each row are in the family sections below and in the list at the end. The B200 per-GPU numbers are computed by dividing NVIDIA's DGX B200 totals (1,440 GB and 64 TB/s) by eight GPUs; B300 is shown as NVIDIA's system total because the per-GPU figure is not stated on the DGX page. FLOPS across rows use different precisions and sparsity settings, which is why the table leads with memory. See TFLOPS and FP4 for how to read the numbers.
NVIDIA Hopper: the workhorse generation
Hopper is still the volume generation. NVIDIA's H100 page lists 80 GB at 3.35 TB/s for the SXM part and 94 GB for the NVL variant. The H200 page lists 141 GB of HBM3e at 4.8 TB/s, which NVIDIA describes as nearly double H100's capacity and 1.4 times its bandwidth.
- NVIDIA H200 guide and H100 vs H200 for the memory and bandwidth step between the two.
- A100 vs H100 and A100 vs V100 cover the Ampere and Volta generations before Hopper; their pages are /gpu/nvidia-a100 and /gpu/nvidia-v100.
- H100, H200: SXM vs NVL vs PCIe explains which form factor to pick.
- NVIDIA Transformer Engine and FP8 covers the number format that made Hopper fast. See also FP8.
- GH200 Grace Hopper pairs a Hopper GPU with a Grace CPU over a 900 GB/s coherent link, per NVIDIA.
- H20 is the region-limited Hopper variant. NVIDIA has not published a datasheet for it, so its page leans on reporting and says so.
Choose Hopper when the model fits and price per hour matters more than speed.
NVIDIA Blackwell and Blackwell Ultra
Blackwell is the current generation for new deployments. DGX B200 has eight GPUs with 1,440 GB of HBM3e and 64 TB/s of memory bandwidth, and 144 PFLOPS sparse FP4 (72 dense), joined by fifth-generation NVLink. DGX B300 uses Blackwell Ultra with 2.1 TB of GPU memory and 144 PFLOPS sparse (108 dense) FP4.
- NVIDIA B200 guide
- NVIDIA B300 Blackwell Ultra guide and B300 vs B200
- NVFP4 vs MXFP4, the 4-bit formats that Blackwell added
- GB200 NVL72, GB300 NVL72 vs GB200 NVL72, and H200 vs B200 vs GB200
- HGX vs DGX vs NVL72 for what the product names mean
At rack scale, NVIDIA's GB200 NVL72 page lists 72 Blackwell GPUs, 13.4 TB of HBM3E, 576 TB/s of GPU memory bandwidth and 720 PFLOPS dense FP4. The GB300 NVL72 page lists 72 Blackwell Ultra GPUs, 20 TB of GPU memory and up to 576 TB/s.
Two desk-side systems use the same architecture: the DGX Spark and the DGX Station GB300. Both are covered with rent-vs-buy math.
NVIDIA Rubin: announced, not the default
NVIDIA's January 5, 2026 announcement states: "NVIDIA Rubin is in full production, and Rubin-based products will be available from partners the second half of 2026" (NVIDIA newsroom). Per NVIDIA's technical blog, the Rubin GPU has up to 288 GB of HBM4, up to 22 TB/s of bandwidth (about 2.8 times Blackwell), and up to 50 PFLOPS of NVFP4. Vera Rubin NVL72 combines 72 Rubin GPUs and 36 Vera CPUs, with 216 to 260 TB/s of rack NVLink (NVIDIA's product page and technical blog differ).
"In production" is not the same as "rentable by the hour." Treat Rubin as a planning target until instances show live prices below.
- NVIDIA Rubin guide
- Vera Rubin NVL72 guide
- Rubin CPX guide: NVIDIA announced it in September 2025 with 128 GB of GDDR7 and 30 PFLOPS of NVFP4, expected at the end of 2026 (NVIDIA newsroom). Later press reporting describes a redesign; NVIDIA has not confirmed it on the pages we checked.
- Rubin vs Blackwell vs Hopper puts all three generations side by side.
AMD Instinct
AMD's line has moved quickly. MI300X carries 192 GB of HBM3 at 5.3 TB/s, MI325X 256 GB of HBM3E at 6 TB/s, and MI355X up to 288 GB of HBM3E at 8 TB/s with FP4 and FP6 support (Tom's Hardware, Next Platform).
- AMD MI355X guide
- AMD MI325X guide
- AMD MI400 and MI450 guide: AMD unveiled the MI455X and Helios rack in July 2026, with 432 GB of HBM4 and 40 PFLOPS of FP4 per GPU as AMD's peak figures, and shipments reported for the end of Q3 2026 ramping through 2027 (StorageReview). We could not verify this against AMD's own page, so treat the timing as reported.
AMD's large memory pools suit models that barely fit on an 80 GB card. Software maturity is the main question; the guides cover ROCm support.
Interconnect and memory
A chip is only as useful as its links. These posts cover the parts that decide scaling:
- What is NVLink, with the NVSwitch and NVLink vs PCIe glossary entries
- InfiniBand vs Ethernet for GPU clusters, plus RDMA and NCCL
- UALink vs NVLink, the open scale-up standard AMD is backing
- HBM3e vs HBM4, the memory generation that moves Rubin and MI455X
Other accelerators
These chips are built by cloud and silicon companies rather than sold as general GPUs.
-
Google TPU Ironwood. Google's documentation lists 192 GB of HBM per chip, about 7.37 TB/s, 4,614 TFLOPs at FP8 and 9,216 chips per pod. See TPU vs GPU.
-
AWS Trainium3. AWS lists 144 GB of HBM3e and 4.9 TB/s per chip, with UltraServers scaling to 144 chips. See Trainium vs NVIDIA.
-
Intel Gaudi 3. 128 GB of HBM2e at 3.7 TB/s per AnandTech's launch coverage. See Gaudi 3 vs NVIDIA.
-
Cerebras WSE-3. A wafer-scale chip with 900,000 cores, 44 GB of on-chip SRAM, 21 PB/s of on-chip bandwidth and a 125 PFLOPS peak claim, per Cerebras and SiliconANGLE's launch coverage. That is a small capacity at extreme bandwidth, a different trade from HBM parts. See Cerebras vs NVIDIA.
-
Etched Sohu. A transformer-only ASIC. Etched's public claims are performance claims without independent benchmarks, and we could not confirm shipping from a primary source. See Etched Sohu vs NVIDIA.
-
Huawei Ascend 950. Per Huawei Connect 2025 coverage, 950PR uses roughly 128 GB of in-house HBM and 950DT 144 GB, with 1 PFLOPS FP8 and 2 PFLOPS FP4; 950DT is slated for Q4 2026. See Ascend vs NVIDIA.
-
NVIDIA Groq 3 LPX. Covered in its own guide; see the Groq 3 LPX guide.
-
Tenstorrent. Compared with NVIDIA in Tenstorrent vs NVIDIA.
-
Qualcomm AI200. Compared with NVIDIA in Qualcomm AI200 vs NVIDIA.
The common thread: each is excellent inside its maker's environment and harder to move workloads onto than a CUDA GPU.
The NVIDIA roadmap, dated only from sources
| Item | What the vendor said | Source |
|---|---|---|
| Rubin | In full production; partner products in the second half of 2026 | NVIDIA, January 5, 2026 |
| Rubin CPX | Expected at the end of 2026 (September 2025 announcement) | NVIDIA |
| AMD MI455X and Helios | Unveiled July 2026; shipments reported from end of Q3 2026 | StorageReview |
| Huawei Ascend 950DT | Q4 2026 | Huawei Connect 2025 |
How to choose
- Start from the model, not the chip. Weights at your precision plus KV cache must fit in memory. A 70B model at 8-bit needs about 70 GB for weights alone (70 billion x 1 byte), so it fits on an H100 with little room; an H200 or MI300X leaves room for context. Our KV cache and quantization entries explain the sizing.
- Decode speed follows bandwidth. Generating tokens reads the weights each step, so higher TB/s means faster output. That is the real gap between H100, H200 and B200.
- Training scale follows interconnect. If a job spans many GPUs, NVLink inside the node and the network outside it matter more than single-chip FLOPS.
- Mixture-of-experts changes the math. Memory capacity decides what loads; only active experts move per token. See mixture-of-experts.
- Pick the oldest chip that fits. If an H100 does the job, a B300 only adds cost. If it does not fit, step up one tier, not three.
- Watch the software. CUDA runs everywhere; ROCm, TPU, Trainium and Gaudi each need porting work.
- Desk or datacenter? For prototyping and privacy, a DGX Spark can be enough; for anything sustained or time-bound, rent. Compare chips directly on our compare pages and the GPU index.
Rent today
Aquanode manages and optimizes GPUs for training and inference workloads, and you can rent the GPUs in the box below on demand. Prices update live, and a row reads "None right now" when no offer is available.
For full price history and specs, see /pricing and the GPU index.
Related reading: who competes with NVIDIA in AI, MI300X vs H100 and H200 for inference, AMD vs NVIDIA GPUs for AI and the best GPUs for AI.
FAQ
What is the best datacenter GPU for AI in 2026?
There is no single best. For capacity and speed, Blackwell (B200, B300) leads what you can rent. For cost-sensitive work where the model fits, H100 and H200 remain widely used. The smallest chip whose memory fits your model is usually the right start.
How much memory does each current NVIDIA datacenter GPU have?
H100 SXM has 80 GB, H200 has 141 GB, a B200 works out to 180 GB per GPU (NVIDIA lists 1,440 GB for eight), and Rubin is specified at up to 288 GB of HBM4.
Is NVIDIA Rubin available?
NVIDIA says Rubin is in full production and partner products arrive in the second half of 2026 (January 5, 2026 announcement). Check the live box above for what can be rented today.
How does AMD compare with NVIDIA?
AMD offers more memory per chip at its tier: 192 GB on MI300X, 256 GB on MI325X, up to 288 GB on MI355X. NVIDIA has the broader software ecosystem. See the AMD guides above.
Are TPUs and Trainium alternatives to GPUs?
Yes, inside Google Cloud and AWS respectively. They are not general rentable GPUs, and workloads usually need porting.
What does HBM have to do with GPU choice?
HBM capacity decides which models fit, and its bandwidth sets decode speed. See HBM and HBM3e vs HBM4.
Sources
- NVIDIA H100
- NVIDIA H200
- NVIDIA Grace Hopper superchip (GH200)
- NVIDIA DGX B200
- NVIDIA DGX B300
- NVIDIA GB200 NVL72
- NVIDIA GB300 NVL72
- NVIDIA DGX Spark
- NVIDIA DGX Station
- NVIDIA Rubin platform announcement, January 5, 2026
- NVIDIA developer blog: Inside the Rubin GPU architecture
- NVIDIA Rubin CPX announcement
- Tom's Hardware: AMD MI355X specs
- Next Platform: AMD MI325X
- StorageReview: AMD MI455X and Helios
- AMD Instinct MI355X datasheet
- AMD Instinct MI325X datasheet
- AMD Instinct MI400 series
- Google Cloud TPU7x (Ironwood) documentation
- AWS Trainium
- AnandTech: Intel Gaudi 3
- Cerebras CS-3
- SiliconANGLE: Cerebras WSE-3 launch, March 13, 2024
- Huawei Connect 2025 keynote