The AMD Instinct MI325X is AMD's 2024 refresh of the MI300X: the same CDNA 3 compute and the same eight-GPU board, with the memory upgraded from 192 GB of HBM3 to 256 GB of HBM3E and from 5.3 TB/s to 6 TB/s. It is the part AMD positioned against NVIDIA's H200, and it is a drop-in swap for an MI300X server.
This guide covers:
- The MI325X spec sheet against the MI300X and the H200
- What actually changed from the MI300X, and what did not
- Which public benchmarks exist, and what they do and do not show
- Power, cooling and software requirements
- When the MI325X is the right choice, and when to look at the MI355X
TL;DR
- It is a memory upgrade. The MI325X keeps the MI300X's 304 compute units and 2.61 PFLOPS of dense FP8, and raises memory to 256 GB of HBM3E at 6 TB/s (AMD datasheet).
- Against the H200, the memory gap is the story. 256 GB against NVIDIA's 141 GB is about 1.8x the capacity. AMD's own press release says the same. Peak dense FP8 is also higher on paper (2,615 against 1,979 TFLOPS).
- It drops into MI300X servers. AMD's datasheet calls it a "drop-in replacement for the Instinct MI300X Platform," on the same Universal Base Board.
- It is not the newest AMD part. The MI355X replaced it as the flagship in 2025, with native FP4 and FP6 support and 288 GB of HBM3E.
- Verdict: pick the MI325X when you want more memory per GPU than an MI300X or H200 offers without moving to a new platform generation. If your model fits comfortably in 192 GB, the MI300X is the simpler choice. If you want FP4, look at the MI355X.
MI325X specs against the MI300X and H200
All AMD figures are peak theoretical numbers from AMD's datasheets. NVIDIA's H200 figures come from NVIDIA's H200 page, which quotes its tensor-core numbers "with sparsity"; the dense value is half of what it prints.
| Spec | MI325X | MI300X | H200 SXM |
|---|---|---|---|
| Architecture | 3rd Gen CDNA | 3rd Gen CDNA | Hopper |
| Compute units | 304 | 304 | not applicable (NVIDIA counts SMs) |
| Peak engine clock | 2.1 GHz | 2.1 GHz | not listed here |
| Memory | 256 GB HBM3E | 192 GB HBM3 | 141 GB HBM3e |
| Memory bandwidth | 6 TB/s | 5.3 TB/s | 4.8 TB/s |
| Dense FP16 / BF16 | 1,307 TFLOPS | 1,307 TFLOPS | 989.5 TFLOPS (computed from 1,979 with sparsity) |
| Dense FP8 | 2,615 TFLOPS | 2,615 TFLOPS | 1,979 TFLOPS (computed from 3,958 with sparsity) |
| FP64 vector | 81.7 TFLOPS | 81.7 TFLOPS | not listed here |
| Maximum board power | 1,000 W | 750 W | up to 700 W (configurable) |
| GPU-to-GPU links | 7 x 128 GB/s Infinity Fabric | 7 x 128 GB/s Infinity Fabric | NVLink (see our NVLink guide) |
| Host interface | PCIe Gen 5 x16 | PCIe Gen 5 x16 | not listed here |
| Eight-GPU memory | 2,048 GB | 1,536 GB | 1,128 GB (computed) |
Three observations from the table:
- The compute columns for the MI325X and MI300X are identical. The upgrade is memory and power. If your workload is compute-bound rather than memory-bound, expect little change from an MI300X.
- The memory gap against the H200 is large. 256 GB against 141 GB is a 1.82x ratio. Bandwidth is 6 TB/s against 4.8 TB/s, a 1.25x ratio (AMD rounds this to 1.3x in its launch release).
- Power went up 33% over the MI300X. The maximum board power is 1,000 W, against 750 W. That is above the H200 SXM's 700 W ceiling.
Architecture and form factor
The MI325X is an OAM module. Per AMD's datasheet it is built on 5nm compute dies and 6nm I/O dies, with eight accelerated compute dies (XCDs) of 38 compute units each, 4 MB of shared L2 cache, and a 256 MB Infinity Cache shared across the XCDs. Memory connects over an 8,192-bit interface, and the memory clock runs up to 6.0 GT/s.
The 288 GB that became 256 GB
AMD first described the MI325X at Computex on June 2, 2024 with 288 GB of HBM3E and 6 TB/s of bandwidth. By the product launch on October 10, 2024, and in the shipping datasheet, the figure is 256 GB. If you read older coverage that quotes 288 GB for the MI325X, it is quoting the pre-launch plan. The 288 GB capacity arrived with the MI350 Series instead. Use 256 GB for the MI325X.
Platform
Eight MI325X modules sit on AMD's Universal Base Board (UBB 2.0) with HGX host connectors, which is why AMD describes the platform as a drop-in for the MI300X. Each GPU has seven Infinity Fabric links to the other seven GPUs, and AMD lists 128 GB/s of bidirectional bandwidth between each pair on the board. A single MI325X can also be partitioned (SR-IOV, up to 8 partitions) for multi-tenant use.
This is a scale-up design for eight GPUs. Beyond the node, you rely on the host network. AMD's rack-scale designs arrive with the later MI400 generation (see our MI400 and MI450 guide).
Performance: what is published
Everything below is AMD's own material or an MLCommons submission published by AMD. We have not benchmarked the MI325X ourselves.
AMD's launch claims
AMD's October 10, 2024 press release says the MI325X delivers "256GB of HBM3E supporting 6.0TB/s offering 1.8X more capacity and 1.3x more bandwidth than the H200," and "1.3X greater peak theoretical FP16 and FP8 compute performance compared to H200." Those are peak specification ratios, and they match the datasheet arithmetic above. The same release gives a latency test on Mistral-7B in FP16 (128 input tokens, 128 output tokens, one GPU each, vLLM on the MI325X). That is a single small-model data point, so we do not generalize from it.
MLPerf Inference v5.0
AMD's ROCm blog says the MLPerf Inference v5.0 results were published on April 2, 2025, and that AMD's first MI325X submissions covered Llama 2 70B (Offline and Server) and Stable Diffusion XL, on an 8-GPU MI325X server. AMD says the MI325X "competes head-to-head with the H200 GPU." The blog shows the comparison as a chart image and gives no tokens-per-second figures in the text, so we are not quoting a number here. To get one, open the MLCommons v5.0 datacenter results and look up the submission IDs AMD lists: 5.0-0001 (paired with 5.0-0060) for Llama 2 70B and 5.0-0002 (also paired with 5.0-0060) for SDXL.
Independent data
Our earlier MI300X vs H100 vs H200 inference post summarizes a May 2025 SemiAnalysis benchmark covering both the MI300X and MI325X. That benchmark found them competitive with or ahead of the H200 at some latency targets on the largest models, and behind at others (H200 won at low-to-medium latency targets on DeepSeek). The conclusion is workload-dependent, which is the honest answer. Read that post for the details and for the software-maturity caveats, and read AMD vs NVIDIA GPUs for AI for the wider picture.
Infrastructure needs
Power. Maximum board power is 1,000 W per module, so a full eight-GPU board is 8 kW before CPUs, NICs and fans, against 6 kW for eight MI300X modules. Check your rack budget before you assume an MI300X slot will take the new part.
Cooling. AMD's datasheet does not specify a cooling method for the MI325X. Ask your server vendor what the specific platform needs at 1,000 W per module.
Networking. Each GPU has a PCIe Gen 5 x16 link (128 GB/s) to the host and uses the same link class for scale-out network bandwidth. AMD's launch release also introduced the Pensando Pollara 400 NIC and Salina DPU for AI networking. See our glossary on RDMA and NCCL.
Software. AMD's datasheet lists PyTorch, TensorFlow and JAX support through ROCm, with the AMD ROCm Developer Hub for containers and documentation. In practice most teams serve on vLLM or SGLang on ROCm. Our ROCm vs CUDA post covers where the stack still trails CUDA.
When to choose the MI325X
Choose the MI325X when:
- You serve a large model where VRAM capacity decides how many GPUs you need. At 256 GB, an eight-GPU board holds 2,048 GB of weights and KV cache.
- You are running long-context or large-batch inference that is memory-bandwidth bound. The 6 TB/s helps there; the unchanged compute does not.
- You already run MI300X servers and want more memory on the same board design.
Choose an MI300X when:
- The model already fits in 192 GB at your serving precision. You keep the lower 750 W budget and the same compute. Our MI300X rental guide covers that part.
Choose an H200 when:
- Your stack is CUDA-first, or you rely on FP8 kernels already tuned for Hopper. See the H200 page and our comparison at /compare/gpu/amd-mi300x-vs-h200.
Choose an MI355X when:
- You want FP4 or FP6 and 288 GB per GPU. It is the current AMD flagship. See the MI355X guide.
Cost: how to turn a benchmark into tokens per GPU-hour
AMD did not publish a tokens-per-second figure for the MI325X in text, so we do not invent one. The method is simple, and you can apply it to any figure you trust:
- Take the Llama 2 70B Server tokens per second for the 8-GPU MI325X submission from MLCommons.
- Divide by 8 to get tokens per second per GPU.
- Multiply by 3,600 to get tokens per GPU-hour.
- Multiply by the live hourly price in the box below, then divide to get a cost per million tokens.
A benchmark with fixed prompts and a fixed model will not match your traffic. Use it to rank hardware, then measure your own workload.
Rent today
Aquanode manages and optimizes GPUs for training and inference workloads, and you can rent the GPUs in the box below on demand. The box shows which of these chips has a live offer right now; a row reading "None right now" means no offer at the moment.
What's next
AMD's June 2025 MI350 Series (MI350X and MI355X) raised memory to 288 GB and added FP4 and FP6. The MI400 Series followed in July 2026 for rack-scale systems. See our MI355X guide, the MI400 and MI450 guide, and the datacenter GPU guide. For memory technology, see HBM3E vs HBM4.
FAQ
How much memory does the MI325X have?
256 GB of HBM3E at 6 TB/s (AMD datasheet). AMD first previewed 288 GB in June 2024, but the shipping part is 256 GB.
What is the difference between the MI325X and the MI300X?
Memory and power. The MI325X has 256 GB of HBM3E at 6 TB/s against 192 GB of HBM3 at 5.3 TB/s, and a 1,000 W ceiling against 750 W. Compute units (304) and dense FP8 throughput (2,615 TFLOPS) are the same.
Is the MI325X better than the H200?
On paper it has 1.8x the memory and 1.3x the peak FP8 and FP16 throughput (AMD's launch release). Real results depend on the workload and latency target. A May 2025 third-party benchmark found it competitive at some latency targets and behind at others, so measure your own model.
Can the MI325X replace an MI300X in an existing server?
AMD's datasheet describes the MI325X platform as a drop-in replacement for the MI300X platform on the same Universal Base Board with HGX host connectors. Check power and cooling headroom first, because the module's ceiling is 1,000 W.
Does the MI325X support FP4?
AMD's MI325X datasheet lists FP8, FP16, BF16, TF32 and INT8 for AI. FP4 and FP6 appear on the later MI350 Series.
Sources
- AMD Instinct MI325X datasheet (specs, power, platform, drop-in claim)
- AMD Instinct MI300X data sheet
- AMD Delivers Leadership AI Performance with AMD Instinct MI325X Accelerators, October 10, 2024 (256 GB, 1.8x and 1.3x claims, shipment timing, Mistral-7B latency test)
- AMD Accelerates Pace of Data Center AI Innovation, Computex, June 2, 2024 (the earlier 288 GB plan)
- AMD Instinct MI325X GPUs in MLPerf Inference v5.0 (AMD ROCm blog)
- NVIDIA H200 Tensor Core GPU (141 GB, 4.8 TB/s, 700 W, sparse tensor figures)
- AMD Instinct MI350 Series launch press release, June 12, 2025
- AMD Instinct MI355X GPU datasheet