GLM-4.7-Flash vs Qwen3-30B-A3B
GLM-4.7-Flash (31.2B parameters) and Qwen3-30B-A3B (30.5B parameters) side by side: the memory each needs at every precision, what it costs to run on a live GPU, and the context window, KV cache and license where they are published. Numbers are computed from the models' published specs; this page does not rank quality.
Side by side
| Fact | GLM-4.7-Flash | Qwen3-30B-A3B |
|---|---|---|
| Parameters | 31.2B (Mixture-of-experts: 4 of 64 experts active per token (exact active-parameter count not stated on the model card)) | 30.5B (~3B active per token (mixture-of-experts; see total parameters above)) |
| Architecture | Multi-head latent attention; mixture of 64 experts, 4 active per token | Grouped-query attention; mixture of 128 experts, 8 active per token |
| Context length | 198K tokens (202,752) | 40K tokens (40,960) |
| License | MIT | Apache 2.0 |
| Published precision | BF16 | BF16 |
| VRAM needed, As published | 69.8 GB | 68.2 GB |
| VRAM needed, FP8 | 34.9 GB | 34.1 GB |
| VRAM needed, INT4 | 17.4 GB | 17.1 GB |
| Cheapest live fit, As published | A100 · $1.21/hr | A100 · $1.21/hr |
| Cheapest live fit, FP8 | L40 · $0.742/hr | L40 · $0.742/hr |
| Cheapest live fit, INT4 | RTX A5000 · $0.176/hr | RTX A5000 · $0.176/hr |
| KV cache per token (16-bit) | 53 KB | 96 KB |
| KV cache at 32k tokens | 1.65 GB | 3.00 GB |
| KV cache at 128k tokens | 6.61 GB | 12.0 GB |
VRAM is the weight size at each precision times a flat 1.2 overhead; see the methodology. The FP8 and INT4 rows need a quantized checkpoint or an engine that quantizes on load. The fit is the lowest-priced single GPU type that holds the model at that precision, or the lowest-priced multi-GPU set (up to 8) when none does. KV cache is for one sequence at 16-bit, computed from each model's config where the attention layout is known.
Which to pick
- Qwen3-30B-A3B needs less VRAM at its published precision (68.2 GB against 69.8 GB), so it fits on a smaller GPU.
- GLM-4.7-Flash lists the longer context window (202,752 tokens against 40,960).
- GLM-4.7-Flash caches less per sequence at 32k tokens (1.7 GB against 3.0 GB), leaving more memory for batching.
- Licenses differ: GLM-4.7-Flash is under MIT and Qwen3-30B-A3B under Apache 2.0. Read both before commercial use.
These follow only from the facts in the table above. Whether either model does your task well is a separate question this page does not answer.
Keep reading
- GLM-4.7-Flash: full VRAM table and live GPU fit
- Qwen3-30B-A3B: full VRAM table and live GPU fit
- The GLM model series
- The Qwen model series
- All models that fit in 80 GB
Other comparisons