AMD MI355X Guide: Specs, MLPerf, vs B200 (2026)

Back
Team Aquanode

Team Aquanode

Sarthak Vaish

Updated OCTOBER 8, 2026Published OCTOBER 8, 2026

The AMD Instinct MI355X is AMD's flagship 2025-generation accelerator: 288 GB of HBM3E, 8 TB/s of memory bandwidth, a 1,400 W power ceiling, and native FP4 and FP6 support on the 4th Gen CDNA architecture. It is the part AMD positions against NVIDIA's B200, and the MI350X is its lower-power sibling with the same memory.

This guide covers:

  • The MI355X spec sheet against the MI350X, MI325X and MI300X
  • What AMD's MLPerf Inference v6.0 submission actually shows
  • Power, cooling and software requirements
  • When to choose it over an MI300X or a B200

TL;DR

  • Memory is the headline. 288 GB of HBM3E per GPU and 8 TB/s, which is 1.5x the capacity of the MI300X's 192 GB and 1.5x its 5.3 TB/s of bandwidth (AMD datasheets).
  • Low-precision compute is the generational jump. AMD lists 10.07 PFLOPS of dense MXFP4 and 5.03 PFLOPS of dense FP8 per GPU, against 2.61 PFLOPS of dense FP8 on the MI300X and MI325X.
  • The price is power. The maximum board power is 1,400 W per module, against 750 W for the MI300X.
  • Independent-style evidence exists. AMD's MLPerf Inference v6.0 submission (April 1, 2026) used the MI355X exclusively, with over one million tokens per second on Llama 2 70B across an 11-node run. The comparisons to NVIDIA in that write-up are AMD's own framing.
  • Verdict: choose the MI355X when you serve large models in FP4 or FP8 and want the most memory per GPU in a current AMD part. Stay on an MI300X if your model already fits in 192 GB and your stack is tuned for it.

MI355X specs against its neighbours

Figures are AMD's peak theoretical numbers from its own datasheets. "Dense" means without structured sparsity. The H100-class comparison lives in our MI300X vs H100 vs H200 post, so this table stays inside the Instinct family.

SpecMI355XMI350XMI325XMI300X
Architecture4th Gen CDNA4th Gen CDNA3rd Gen CDNA3rd Gen CDNA
Compute units256256304304
Peak engine clock2.4 GHz2.2 GHz2.1 GHz2.1 GHz
Memory288 GB HBM3E288 GB HBM3E256 GB HBM3E192 GB HBM3
Memory bandwidth8 TB/s8 TB/s6 TB/s5.3 TB/s
Dense FP16 / BF162.52 PFLOPS2.31 PFLOPS1.31 PFLOPS1.31 PFLOPS
Dense FP85.03 PFLOPS4.61 PFLOPS2.61 PFLOPS2.61 PFLOPS
Dense MXFP6 / MXFP410.07 / 10.07 PFLOPS9.23 / 9.23 PFLOPSnot listednot listed
Maximum board power1,400 W1,000 W1,000 W750 W
Infinity Fabric links7 x 153.6 GB/s7 x 153.6 GB/s7 x 128 GB/s7 x 128 GB/s
Host interfacePCIe Gen 5 x16PCIe Gen 5 x16PCIe Gen 5 x16PCIe Gen 5 x16

A few readings of that table:

  • The MI355X and MI350X share memory and compute-unit count. The MI355X buys roughly 9% more peak FP8 and FP4 throughput by running a higher clock (2.4 GHz against 2.2 GHz) at 400 W more power. That is AMD's own framing: the datasheet says the extra power lets the MI355X "sustain higher performance over time, minimizing throttling."
  • Peak FP8 per GPU roughly doubles from the MI325X to the MI355X (2.61 to 5.03 PFLOPS dense), even though the compute-unit count drops from 304 to 256.
  • The MI325X and MI300X datasheets list no MXFP6 or MXFP4 rows; AMD's MI355X datasheet presents them as new. If your model is quantized to FP4, the MI355X is the Instinct generation AMD built for it. See our glossary entries on FP4, FP8 and quantization.

Architecture and form factor

The MI355X is an OAM module built from 3nm and 6nm TSMC dies. Per AMD's datasheet, each module contains eight accelerated compute dies (XCDs), each with 32 compute units, 32 KB of L1 cache per compute unit, 4 MB of shared L2, and a shared 256 MB AMD Infinity Cache. Two mirrored I/O dies tie the XCDs to eight stacks of HBM3E over an 8,192-bit interface.

Three architectural details matter for deployment:

  • Eight-GPU platform. Eight MI355X modules sit on AMD's Universal Base Board (UBB 2.0), which the datasheet says "can be integrated into server designs as compact as 2U." Each GPU reaches the other seven over Infinity Fabric. AMD lists seven links of 153.6 GB/s each per GPU, and 160 GB/s of bidirectional bandwidth between each pair of GPUs on the board. Eight GPUs at 288 GB each give you 2,304 GB of HBM3E in one platform.
  • Partitioning. A single MI355X can be split into up to eight partitions of 36 GB each (SR-IOV), and the memory can be presented as one or four partitions. That matters if you want to run several small models per GPU.
  • Scale-out. Beyond the eight-GPU board, you use the host network. The datasheet states that the MI355X "supports massive Ethernet-based AI networking," and AMD's rack design pairs it with Pensando Pollara NICs. There is no NVLink-style switched fabric across nodes here; for that you need the rack-scale parts covered in our MI400 and MI450 guide.

There is also a PCIe sibling. In July 2026 AMD launched the MI350P, a PCIe add-in card with 144 GB of HBM3E and a 600 W maximum board power (450 W configurable). It is a different product from the MI355X OAM module, aimed at fitting into existing servers.

Performance: what AMD and MLPerf publish

All figures in this section are AMD's or MLCommons-submission figures published by AMD. We did not run them.

MLPerf Inference v6.0

AMD's ROCm blog states that the MLPerf Inference v6.0 results "were released on April 1st 2026" and that AMD's submissions used the MI355X. Its results table lists tokens per second as follows.

BenchmarkNodesOfflineServerInteractive
Llama 2 70B1103,480100,28273,608
gpt-oss-120b195,00482,136not run
Llama 2 70B11 (87 GPUs)1,042,1101,016,380785,522
gpt-oss-120b121,031,070900,054not run

AMD's write-up also characterizes the comparison with NVIDIA. For Llama 2 70B, the MI355X "tied with Nvidia in Offline mode" and was competitive in Server mode. Against the B300, AMD reports 92% of NVIDIA's offline result, 93% of server and 104% of interactive. For gpt-oss-120b, AMD reports up to 115% in server mode and 111% in offline mode against the two best results from NVIDIA's OEM partners, because NVIDIA itself did not submit a B200 score for that benchmark. Treat all of those as AMD's reading of the public MLCommons tables, and check the tables before you quote a ratio. AMD also flags its Wan2.2 scores as unofficial because they were captured while the submissions were still under review.

Two things worth noticing. First, scaling held up: 87 GPUs delivered about 10x the single-node Llama 2 70B throughput. Second, the interactive scenario, which has the tightest latency target, is where AMD reports its best ratio against the B300.

Launch claims

At Advancing AI on June 12, 2025, AMD described the MI350 Series (MI350X and MI355X) as delivering "a 4x, generation-on-generation AI compute increase and a 35x generational leap in inferencing." It also claimed up to 40% more tokens per dollar than competing solutions. The footnote for that second claim rests on a Llama 3.1 405B FP4 test against published B200 results, plus current B200 cloud pricing and expected MI355X instance pricing. That is a pricing assumption, not a measurement, so we do not carry the percentage forward.

Infrastructure needs

Power. The maximum board power is 1,400 W per module. Eight modules is 11.2 kW for the GPUs alone, before CPUs, NICs and fans. The MI300X at 750 W was 6 kW for the same eight GPUs.

Cooling. AMD's MI350 datasheet describes both air-cooled and direct-liquid-cooled options across the series. At 1,400 W, confirm with the server vendor which cooling your specific MI355X platform requires. The MI350X at 1,000 W is the lower-power option for air-cooled halls.

Networking. Each GPU has a PCIe Gen 5 x16 link (128 GB/s) to the host. AMD recommends Ethernet-based scale-out, with Pensando Pollara 400G NICs in its reference designs. Plan your cluster fabric for RoCE or Ultra Ethernet rather than assuming an InfiniBand-only design. See our glossary on RDMA and NCCL for the collective-communication side.

Software. The MI355X is supported by AMD's ROCm stack. AMD's datasheet lists PyTorch, TensorFlow, JAX, ONNX Runtime, SGLang, Triton and vLLM, and notes "Day-0 support" for optimized models. The AMD GPU Operator handles Kubernetes deployment. For an honest view of where ROCm still lags CUDA, read our ROCm vs CUDA comparison and the AMD vs NVIDIA GPUs for AI post.

When to choose the MI355X

Choose the MI355X when:

  • You serve models that benefit from 288 GB per GPU: 70B to 400B-class models where fewer tensor-parallel splits means less communication overhead, or long-context serving where the KV cache dominates memory.
  • Your model is, or can be, quantized to FP8 or FP4. The MI355X's gain over the MI300X is concentrated in those datatypes.
  • You run a mixture-of-experts model and want the memory to hold more experts per GPU. See mixture of experts.
  • You are building on an open software stack and want a second silicon supplier.

Stay on an MI300X when:

  • Your model already fits in 192 GB at the precision you serve, and your stack is validated on it. The jump to the MI355X buys you headroom, not necessarily a faster result.
  • Your facility cannot supply 1,400 W modules.

Choose a B200 when:

  • Your serving stack is CUDA-first and you do not want to own ROCm kernel work. The B200 side of this comparison is covered in our B200 guide.

For platform-level context, here is how the eight-GPU systems compare on the vendors' own peak numbers. NVIDIA's HGX page lists 144 PFLOPS of FP4 with sparsity for HGX B200, which its footnote halves to 72 PFLOPS dense, and 72 PFLOPS of FP8 with sparsity (36 PFLOPS dense by the same halving). AMD's per-GPU datasheet figures, multiplied by eight, give about 80.5 PFLOPS of dense MXFP4 and 40.3 PFLOPS of dense FP8 for an MI355X platform. Memory is 2.3 TB of HBM3E against 1.4 TB for the HGX B200. These are peak theoretical figures from two vendors with different footnote conventions, so use them to size memory, not to predict your throughput. For the NVIDIA side of the architecture, see B300 vs B200.

Cost: tokens per GPU-hour

We do not publish an hourly price here. Use the throughput numbers above and the live price in the box below.

From AMD's MLPerf table (computed, not measured by us):

  • Llama 2 70B Server, one node: 100,282 tokens per second, or about 361 million tokens per node-hour. An MI355X platform has eight GPUs, so that is about 12,500 tokens per second per GPU, or roughly 45 million tokens per GPU-hour, assuming the submission used the standard eight-GPU node.
  • Llama 2 70B Server, 11 nodes and 87 GPUs: 1,016,380 tokens per second, which is about 11,700 tokens per second per GPU, or roughly 42 million tokens per GPU-hour. That is the figure to use for a multi-node deployment, because it includes the scaling loss.
  • gpt-oss-120b Server, one node: 82,136 tokens per second, about 10,300 per GPU, or roughly 37 million tokens per GPU-hour.

Multiply the tokens per GPU-hour for your workload by the live hourly price below and you have a cost per million tokens. MLPerf is a controlled benchmark with fixed models, so your own prompts, batch sizes and latency target will land somewhere else. Treat these as an upper-bound shape, not a quote.

Rent today

Aquanode manages and optimizes GPUs for training and inference workloads, and you can rent the GPUs in the box below on demand. If the MI355X shows "None right now," the MI300X is the nearest AMD part, and the B200 is the nearest NVIDIA part.

What's next

AMD's next generation is the MI400 Series. AMD launched it on July 23, 2026, with the MI455X for the Helios rack and the MI430X for HPC and sovereign AI, and says volume deployments of Helios are expected in the second half of 2026. Those are rack-scale parts, not drop-in replacements for an eight-GPU MI355X server. Details and status are in our MI400 and MI450 guide. For the wider datacenter picture, see the datacenter GPU guide, and for memory technology see HBM3E vs HBM4.

FAQ

What is the difference between the MI355X and the MI350X?

Both have 288 GB of HBM3E, 8 TB/s of bandwidth and 256 compute units. The MI355X runs at a 2.4 GHz peak clock and 1,400 W maximum board power. The MI350X runs at 2.2 GHz and 1,000 W, with dense FP8 of 4.61 PFLOPS against 5.03 PFLOPS (AMD datasheets).

How much memory does the MI355X have?

288 GB of HBM3E at 8 TB/s per GPU. An eight-GPU platform therefore has 2,304 GB. The MI300X has 192 GB and the MI325X has 256 GB.

Is the MI355X faster than the B200?

It depends on workload, precision and latency target. AMD says its MLPerf Inference v6.0 Llama 2 70B result tied NVIDIA in the offline scenario and was competitive in server, and that it exceeded NVIDIA's OEM-submitted results on gpt-oss-120b. Those are AMD's readings of the public tables. Check the MLCommons results for the benchmark and scenario that matches your serving setup.

Does the MI355X support FP4?

Yes. AMD lists MXFP4 and MXFP6 at 10.07 PFLOPS dense per GPU. The MI300X and MI325X datasheets list no FP4 or FP6 rows.

What software runs on the MI355X?

ROCm, with PyTorch, TensorFlow, JAX, ONNX Runtime, SGLang, Triton and vLLM listed in AMD's datasheet. AMD also provides a GPU Operator for Kubernetes.

Is the MI355X liquid cooled?

AMD's MI350 datasheet lists both air-cooled and direct-liquid-cooled options for the series, and the MI355X has a 1,400 W ceiling. Confirm the cooling requirement with your server vendor for the exact platform.

Sources

#datacenter gpu#amd instinct#mi355x#mi350x#cdna 4#rocm#mlperf

Submit the job. Everything after that is ours.

Sign up in 60 seconds. Pay for the GPU minutes you actually use.

© 2026 Aquanode. All rights reserved.

All trademarks, logos and brand names are the property of their respective owners.