Hy3-preview vs MiniMax-M2

Hy3-preview (298.8B parameters) and MiniMax-M2 (228.7B parameters) side by side: the memory each needs at every precision, what it costs to run on a live GPU, and the context window, KV cache and license where they are published. Numbers are computed from the models' published specs; this page does not rank quality.

Side by side

FactHy3-previewMiniMax-M2
Parameters298.8B228.7B (Mixture-of-experts: 8 of 256 experts active per token (10B active parameters, stated on the model card))
ArchitectureGrouped-query attention; mixture of 192 experts, 8 active per tokenGrouped-query attention; mixture of 256 experts, 8 active per token
Context length262,144 tokens192K tokens (196,608)
License–Modified MIT
Published precisionBF16F8_E4M3
VRAM needed, As published668 GB256 GB
VRAM needed, FP8334 GBNot a smaller option
VRAM needed, INT4167 GB128 GB
Cheapest live fit, As publishedRTX PRO 6000 × 7 · $10.29/hrRTX 4080 Super × 8 · $2.71/hr
Cheapest live fit, FP8L40 × 7 · $5.31/hr–
Cheapest live fit, INT4RTX A5000 × 7 · $1.23/hrRTX A5000 × 6 · $1.06/hr
KV cache per token (16-bit)320 KB248 KB
KV cache at 32k tokens10.0 GB7.75 GB
KV cache at 128k tokens40.0 GB31.0 GB

VRAM is the weight size at each precision times a flat 1.2 overhead; see the methodology. The FP8 and INT4 rows need a quantized checkpoint or an engine that quantizes on load. The fit is the lowest-priced single GPU type that holds the model at that precision, or the lowest-priced multi-GPU set (up to 8) when none does. KV cache is for one sequence at 16-bit, computed from each model's config where the attention layout is known.

Which to pick

  • MiniMax-M2 needs less VRAM at its published precision (256 GB against 668 GB), so it fits on a smaller GPU.
  • Hy3-preview lists the longer context window (262,144 tokens against 196,608).
  • MiniMax-M2 caches less per sequence at 32k tokens (7.8 GB against 10.0 GB), leaving more memory for batching.
  • MiniMax-M2 has the cheaper live GPU fit at its published precision ($2.71/hr against $10.29/hr).

These follow only from the facts in the table above. Whether either model does your task well is a separate question this page does not answer.

Keep reading

Other comparisons

Submit the job. Everything after that is ours.

Sign up in 60 seconds. Pay for the GPU minutes you actually use.

© 2026 Aquanode. All rights reserved.

All trademarks, logos and brand names are the property of their respective owners.