AMD MI325X Guide: Specs, 256GB HBM3E, vs H200 (2026)

Back
Team Aquanode

Team Aquanode

Sarthak Vaish

Updated OCTOBER 8, 2026Published OCTOBER 8, 2026

The AMD Instinct MI325X is AMD's 2024 refresh of the MI300X: the same CDNA 3 compute and the same eight-GPU board, with the memory upgraded from 192 GB of HBM3 to 256 GB of HBM3E and from 5.3 TB/s to 6 TB/s. It is the part AMD positioned against NVIDIA's H200, and it is a drop-in swap for an MI300X server.

This guide covers:

  • The MI325X spec sheet against the MI300X and the H200
  • What actually changed from the MI300X, and what did not
  • Which public benchmarks exist, and what they do and do not show
  • Power, cooling and software requirements
  • When the MI325X is the right choice, and when to look at the MI355X

TL;DR

  • It is a memory upgrade. The MI325X keeps the MI300X's 304 compute units and 2.61 PFLOPS of dense FP8, and raises memory to 256 GB of HBM3E at 6 TB/s (AMD datasheet).
  • Against the H200, the memory gap is the story. 256 GB against NVIDIA's 141 GB is about 1.8x the capacity. AMD's own press release says the same. Peak dense FP8 is also higher on paper (2,615 against 1,979 TFLOPS).
  • It drops into MI300X servers. AMD's datasheet calls it a "drop-in replacement for the Instinct MI300X Platform," on the same Universal Base Board.
  • It is not the newest AMD part. The MI355X replaced it as the flagship in 2025, with native FP4 and FP6 support and 288 GB of HBM3E.
  • Verdict: pick the MI325X when you want more memory per GPU than an MI300X or H200 offers without moving to a new platform generation. If your model fits comfortably in 192 GB, the MI300X is the simpler choice. If you want FP4, look at the MI355X.

MI325X specs against the MI300X and H200

All AMD figures are peak theoretical numbers from AMD's datasheets. NVIDIA's H200 figures come from NVIDIA's H200 page, which quotes its tensor-core numbers "with sparsity"; the dense value is half of what it prints.

SpecMI325XMI300XH200 SXM
Architecture3rd Gen CDNA3rd Gen CDNAHopper
Compute units304304not applicable (NVIDIA counts SMs)
Peak engine clock2.1 GHz2.1 GHznot listed here
Memory256 GB HBM3E192 GB HBM3141 GB HBM3e
Memory bandwidth6 TB/s5.3 TB/s4.8 TB/s
Dense FP16 / BF161,307 TFLOPS1,307 TFLOPS989.5 TFLOPS (computed from 1,979 with sparsity)
Dense FP82,615 TFLOPS2,615 TFLOPS1,979 TFLOPS (computed from 3,958 with sparsity)
FP64 vector81.7 TFLOPS81.7 TFLOPSnot listed here
Maximum board power1,000 W750 Wup to 700 W (configurable)
GPU-to-GPU links7 x 128 GB/s Infinity Fabric7 x 128 GB/s Infinity FabricNVLink (see our NVLink guide)
Host interfacePCIe Gen 5 x16PCIe Gen 5 x16not listed here
Eight-GPU memory2,048 GB1,536 GB1,128 GB (computed)

Three observations from the table:

  • The compute columns for the MI325X and MI300X are identical. The upgrade is memory and power. If your workload is compute-bound rather than memory-bound, expect little change from an MI300X.
  • The memory gap against the H200 is large. 256 GB against 141 GB is a 1.82x ratio. Bandwidth is 6 TB/s against 4.8 TB/s, a 1.25x ratio (AMD rounds this to 1.3x in its launch release).
  • Power went up 33% over the MI300X. The maximum board power is 1,000 W, against 750 W. That is above the H200 SXM's 700 W ceiling.

Architecture and form factor

The MI325X is an OAM module. Per AMD's datasheet it is built on 5nm compute dies and 6nm I/O dies, with eight accelerated compute dies (XCDs) of 38 compute units each, 4 MB of shared L2 cache, and a 256 MB Infinity Cache shared across the XCDs. Memory connects over an 8,192-bit interface, and the memory clock runs up to 6.0 GT/s.

The 288 GB that became 256 GB

AMD first described the MI325X at Computex on June 2, 2024 with 288 GB of HBM3E and 6 TB/s of bandwidth. By the product launch on October 10, 2024, and in the shipping datasheet, the figure is 256 GB. If you read older coverage that quotes 288 GB for the MI325X, it is quoting the pre-launch plan. The 288 GB capacity arrived with the MI350 Series instead. Use 256 GB for the MI325X.

Platform

Eight MI325X modules sit on AMD's Universal Base Board (UBB 2.0) with HGX host connectors, which is why AMD describes the platform as a drop-in for the MI300X. Each GPU has seven Infinity Fabric links to the other seven GPUs, and AMD lists 128 GB/s of bidirectional bandwidth between each pair on the board. A single MI325X can also be partitioned (SR-IOV, up to 8 partitions) for multi-tenant use.

This is a scale-up design for eight GPUs. Beyond the node, you rely on the host network. AMD's rack-scale designs arrive with the later MI400 generation (see our MI400 and MI450 guide).

Performance: what is published

Everything below is AMD's own material or an MLCommons submission published by AMD. We have not benchmarked the MI325X ourselves.

AMD's launch claims

AMD's October 10, 2024 press release says the MI325X delivers "256GB of HBM3E supporting 6.0TB/s offering 1.8X more capacity and 1.3x more bandwidth than the H200," and "1.3X greater peak theoretical FP16 and FP8 compute performance compared to H200." Those are peak specification ratios, and they match the datasheet arithmetic above. The same release gives a latency test on Mistral-7B in FP16 (128 input tokens, 128 output tokens, one GPU each, vLLM on the MI325X). That is a single small-model data point, so we do not generalize from it.

MLPerf Inference v5.0

AMD's ROCm blog says the MLPerf Inference v5.0 results were published on April 2, 2025, and that AMD's first MI325X submissions covered Llama 2 70B (Offline and Server) and Stable Diffusion XL, on an 8-GPU MI325X server. AMD says the MI325X "competes head-to-head with the H200 GPU." The blog shows the comparison as a chart image and gives no tokens-per-second figures in the text, so we are not quoting a number here. To get one, open the MLCommons v5.0 datacenter results and look up the submission IDs AMD lists: 5.0-0001 (paired with 5.0-0060) for Llama 2 70B and 5.0-0002 (also paired with 5.0-0060) for SDXL.

Independent data

Our earlier MI300X vs H100 vs H200 inference post summarizes a May 2025 SemiAnalysis benchmark covering both the MI300X and MI325X. That benchmark found them competitive with or ahead of the H200 at some latency targets on the largest models, and behind at others (H200 won at low-to-medium latency targets on DeepSeek). The conclusion is workload-dependent, which is the honest answer. Read that post for the details and for the software-maturity caveats, and read AMD vs NVIDIA GPUs for AI for the wider picture.

Infrastructure needs

Power. Maximum board power is 1,000 W per module, so a full eight-GPU board is 8 kW before CPUs, NICs and fans, against 6 kW for eight MI300X modules. Check your rack budget before you assume an MI300X slot will take the new part.

Cooling. AMD's datasheet does not specify a cooling method for the MI325X. Ask your server vendor what the specific platform needs at 1,000 W per module.

Networking. Each GPU has a PCIe Gen 5 x16 link (128 GB/s) to the host and uses the same link class for scale-out network bandwidth. AMD's launch release also introduced the Pensando Pollara 400 NIC and Salina DPU for AI networking. See our glossary on RDMA and NCCL.

Software. AMD's datasheet lists PyTorch, TensorFlow and JAX support through ROCm, with the AMD ROCm Developer Hub for containers and documentation. In practice most teams serve on vLLM or SGLang on ROCm. Our ROCm vs CUDA post covers where the stack still trails CUDA.

When to choose the MI325X

Choose the MI325X when:

  • You serve a large model where VRAM capacity decides how many GPUs you need. At 256 GB, an eight-GPU board holds 2,048 GB of weights and KV cache.
  • You are running long-context or large-batch inference that is memory-bandwidth bound. The 6 TB/s helps there; the unchanged compute does not.
  • You already run MI300X servers and want more memory on the same board design.

Choose an MI300X when:

  • The model already fits in 192 GB at your serving precision. You keep the lower 750 W budget and the same compute. Our MI300X rental guide covers that part.

Choose an H200 when:

Choose an MI355X when:

  • You want FP4 or FP6 and 288 GB per GPU. It is the current AMD flagship. See the MI355X guide.

Cost: how to turn a benchmark into tokens per GPU-hour

AMD did not publish a tokens-per-second figure for the MI325X in text, so we do not invent one. The method is simple, and you can apply it to any figure you trust:

  1. Take the Llama 2 70B Server tokens per second for the 8-GPU MI325X submission from MLCommons.
  2. Divide by 8 to get tokens per second per GPU.
  3. Multiply by 3,600 to get tokens per GPU-hour.
  4. Multiply by the live hourly price in the box below, then divide to get a cost per million tokens.

A benchmark with fixed prompts and a fixed model will not match your traffic. Use it to rank hardware, then measure your own workload.

Rent today

Aquanode manages and optimizes GPUs for training and inference workloads, and you can rent the GPUs in the box below on demand. The box shows which of these chips has a live offer right now; a row reading "None right now" means no offer at the moment.

What's next

AMD's June 2025 MI350 Series (MI350X and MI355X) raised memory to 288 GB and added FP4 and FP6. The MI400 Series followed in July 2026 for rack-scale systems. See our MI355X guide, the MI400 and MI450 guide, and the datacenter GPU guide. For memory technology, see HBM3E vs HBM4.

FAQ

How much memory does the MI325X have?

256 GB of HBM3E at 6 TB/s (AMD datasheet). AMD first previewed 288 GB in June 2024, but the shipping part is 256 GB.

What is the difference between the MI325X and the MI300X?

Memory and power. The MI325X has 256 GB of HBM3E at 6 TB/s against 192 GB of HBM3 at 5.3 TB/s, and a 1,000 W ceiling against 750 W. Compute units (304) and dense FP8 throughput (2,615 TFLOPS) are the same.

Is the MI325X better than the H200?

On paper it has 1.8x the memory and 1.3x the peak FP8 and FP16 throughput (AMD's launch release). Real results depend on the workload and latency target. A May 2025 third-party benchmark found it competitive at some latency targets and behind at others, so measure your own model.

Can the MI325X replace an MI300X in an existing server?

AMD's datasheet describes the MI325X platform as a drop-in replacement for the MI300X platform on the same Universal Base Board with HGX host connectors. Check power and cooling headroom first, because the module's ceiling is 1,000 W.

Does the MI325X support FP4?

AMD's MI325X datasheet lists FP8, FP16, BF16, TF32 and INT8 for AI. FP4 and FP6 appear on the later MI350 Series.

Sources

#datacenter gpu#amd instinct#mi325x#mi300x#cdna 3#hbm3e#h200

Submit the job. Everything after that is ours.

Sign up in 60 seconds. Pay for the GPU minutes you actually use.

© 2026 Aquanode. All rights reserved.

All trademarks, logos and brand names are the property of their respective owners.