AMD MI355X GPU: Specs, VRAM, Price & Benchmarks (2026)

The AMD MI355X is a real GPU (launched 2025), but no provider on Aquanode is listing it for rental right now, so there is no live hourly price to quote. Availability changes as providers add and retire hardware; the specs, VRAM and price context below still apply.

How much VRAM does the AMD MI355X have?

The AMD MI355X has 288GB HBM3E, with 8 TB/s of peak memory bandwidth.

AMD MI355X VRAM calculator: check which models fit in its memory at each precision.

All models that fit in 288 GB: the open models whose weights and overhead fit, at native, FP8 and INT4 precision.

What can the AMD MI355X run?

Popular open models from small to frontier scale, with the memory each needs and how many AMD MI355X cards (288GB HBM3E each) that takes.

ModelAs publishedFP8INT4
Qwen/Qwen3-8B 8.2BBF16: ~18.3 GB, 1 GPUFP8: ~9.2 GB, 1 GPUINT4: ~4.6 GB, 1 GPU
Qwen/Qwen2.5-14B-Instruct 14.8BBF16: ~33 GB, 1 GPUFP8: ~16.5 GB, 1 GPUINT4: ~8.3 GB, 1 GPU
Qwen/Qwen3-32B 32.8BBF16: ~73.2 GB, 1 GPUFP8: ~36.6 GB, 1 GPUINT4: ~18.3 GB, 1 GPU
Qwen/Qwen-72B 72.3BBF16: ~162 GB, 1 GPUFP8: ~80.8 GB, 1 GPUINT4: ~40.4 GB, 1 GPU
MiniMaxAI/MiniMax-M2.7 228.7BFP8: ~256 GB, 1 GPU–INT4: ~128 GB, 1 GPU
deepseek-ai/DeepSeek-R1 684.5BFP8: ~765 GB, 3 GPUs–INT4: ~383 GB, 2 GPUs

Estimates: weights at the stated precision plus a flat 20% for KV cache and overhead, at a moderate context length. A dash means the precision is not offered for that model (it is already published at that size). INT4 needs a published quantized checkpoint. Open any model for a per-GPU breakdown, or use the AMD MI355X VRAM calculator.

AMD MI355X specs

Architecturelaunched 2025
VRAM288GB HBM3E
Memory bandwidth8 TB/s
FP16 / BF16 tensor throughput2,500 TFLOPS (peak, dense)
FP8 tensor throughput5,000 TFLOPS (peak, dense)
InterconnectInfinity Fabric, 7 links at 153 GB/s peak each
TDP1400W typical board power
Form factorOAM

Specs sourced from the vendor's public product page. See the source.

GPU Glossary: What is VRAM?, HBM, Tensor Cores, CUDA Cores, TFLOPS, NVLink vs PCIe

Good for

288GB on a single card, as much memory per GPU as any part in this guide, at 8 TB/s of bandwidth. A 70B model fits in FP16 with room for a long-context KV cache, and a 405B model fits in INT4 on one card. Suited to memory-bound inference of very large models and to training runs where fewer, larger GPUs simplify parallelism.

Not good for

Runs on AMD's ROCm software stack rather than CUDA, so some CUDA-only tooling and custom kernels need a ROCm port or do not run at all. At 1400W it also needs a platform built for it, and it is not a part you will find in a PCIe workstation.

Related guides

Other models in the same generation, or see how it compares in the GPU benchmarks and specs table. The full list is in the GPU index.

Related reading: AMD MI355X guide.

Get notified when the price drops

GPU supply moves hourly. Tell us what you're waiting for and we'll email you when a matching offer appears across any provider we track.

One email per matching alert. Unsubscribe any time.

Submit the job. Everything after that is ours.

Sign up in 60 seconds. Pay for the GPU minutes you actually use.

© 2026 Aquanode. All rights reserved.

All trademarks, logos and brand names are the property of their respective owners.