What is Mixture of Experts (MoE)?

Abbreviated MoE

A Mixture of Experts (MoE) model is a neural network whose feed-forward layers are split into many separate sub-networks called experts, with a small router that sends each token to only a few of them. The model has a large total parameter count, but only a fraction of those parameters, the active parameters, do any work for a given token.

The consequence for anyone renting GPUs: MoE saves compute, not memory. Every expert's weights must stay resident in VRAM, because the router can pick any of them for the next token.

How routing works

In a dense transformer, every token passes through the same feed-forward block in each layer. In an MoE layer, that block is replaced by several experts (each its own feed-forward network) plus a router. The router scores the experts for each token, picks the top few, and combines their outputs weighted by those scores. The attention layers are typically shared by all tokens rather than split into experts. Different tokens in the same sequence can go to different experts, and each layer makes its own choice.

Mixtral 8x7B from Mistral AI is the standard example. Each layer has 8 experts and the router picks 2 per token. The model has about 47B total parameters and about 13B active per token. It is not 8 x 7B = 56B because the attention layers and embeddings are shared across experts.

Total versus active parameters

Total parameters set how much VRAM the weights need. Active parameters set how much arithmetic each token costs, at roughly 2 floating-point operations per active parameter per token (a standard rule of thumb). For Mixtral that is about 2 x 13B = 26 billion operations per token, against 2 x 47B = 94 billion if every parameter were active.

There is a catch on memory bandwidth. With a single request, only a few experts are read per step. With a large batch, different tokens tend to pick different experts, so most of the weights get read on every step and the bandwidth saving shrinks. The compute saving remains.

What it means when you pick a GPU

Size VRAM by total parameters and compute by active parameters. Here is Mixtral 8x7B's weight memory:

PrecisionComputationWeights
16-bit47B x 2 bytesabout 94 GB
8-bit47B x 1 byteabout 47 GB
4-bit47B x 0.5 bytesabout 23.5 GB

At 16-bit, 94 GB does not fit on a single H100 (80GB HBM3). An H200 (141GB HBM3e) holds it with about 47 GB left over (141 minus 94) for the KV cache and overhead, and a B200 (180GB) or an AMD MI300X (192GB) has more room still. Quantization changes the picture: at 8-bit the weights fit on an H100 with space for a modest cache, and at 4-bit they approach the full 24GB of an RTX 4090, leaving essentially nothing for the cache.

In other words, the model computes like a roughly 13B dense model but needs the memory of a roughly 47B one. Paying for a card with a lot of VRAM and modest compute can be a better match than a card with the most compute.

If the weights do not fit on one GPU, they are split across several. The MoE-specific approach is expert parallelism, where different experts live on different GPUs and tokens are sent between GPUs to reach their experts in every layer. That makes inter-GPU communication, typically through a collective library such as NCCL, part of the cost per token. If the weights fit on one card, you avoid that traffic entirely.

Not sure which card fits your model? The GPU recommender can narrow it down, and Aquanode rents GPUs by the hour.

Building on GPUs? Aquanode runs the workload.

Deploy on H100, H200, B200, A100 and MI300X across a multi-provider marketplace, without racking your own hardware or committing to one cloud's spec sheet.

See also

Submit the job. Everything after that is ours.

Sign up in 60 seconds. Pay for the GPU minutes you actually use.

© 2026 Aquanode. All rights reserved.

All trademarks, logos and brand names are the property of their respective owners.