What is Mixture of Experts (MoE)?
Abbreviated MoE
A Mixture of Experts (MoE) model is a neural network whose feed-forward layers are split into many separate sub-networks called experts, with a small router that sends each token to only a few of them. The model has a large total parameter count, but only a fraction of those parameters, the active parameters, do any work for a given token.
The consequence for anyone renting GPUs: MoE saves compute, not memory. Every expert's weights must stay resident in VRAM, because the router can pick any of them for the next token.
How routing works
In a dense transformer, every token passes through the same feed-forward block in each layer. In an MoE layer, that block is replaced by several experts (each its own feed-forward network) plus a router. The router scores the experts for each token, picks the top few, and combines their outputs weighted by those scores. The attention layers are typically shared by all tokens rather than split into experts. Different tokens in the same sequence can go to different experts, and each layer makes its own choice.
Mixtral 8x7B from Mistral AI is the standard example. Each layer has 8 experts and the router picks 2 per token. The model has about 47B total parameters and about 13B active per token. It is not 8 x 7B = 56B because the attention layers and embeddings are shared across experts.
Total versus active parameters
Total parameters set how much VRAM the weights need. Active parameters set how much arithmetic each token costs, at roughly 2 floating-point operations per active parameter per token (a standard rule of thumb). For Mixtral that is about 2 x 13B = 26 billion operations per token, against 2 x 47B = 94 billion if every parameter were active.
There is a catch on memory bandwidth. With a single request, only a few experts are read per step. With a large batch, different tokens tend to pick different experts, so most of the weights get read on every step and the bandwidth saving shrinks. The compute saving remains.
What it means when you pick a GPU
Size VRAM by total parameters and compute by active parameters. Here is Mixtral 8x7B's weight memory:
| Precision | Computation | Weights |
|---|---|---|
| 16-bit | 47B x 2 bytes | about 94 GB |
| 8-bit | 47B x 1 byte | about 47 GB |
| 4-bit | 47B x 0.5 bytes | about 23.5 GB |
At 16-bit, 94 GB does not fit on a single H100 (80GB HBM3). An H200 (141GB HBM3e) holds it with about 47 GB left over (141 minus 94) for the KV cache and overhead, and a B200 (180GB) or an AMD MI300X (192GB) has more room still. Quantization changes the picture: at 8-bit the weights fit on an H100 with space for a modest cache, and at 4-bit they approach the full 24GB of an RTX 4090, leaving essentially nothing for the cache.
In other words, the model computes like a roughly 13B dense model but needs the memory of a roughly 47B one. Paying for a card with a lot of VRAM and modest compute can be a better match than a card with the most compute.
If the weights do not fit on one GPU, they are split across several. The MoE-specific approach is expert parallelism, where different experts live on different GPUs and tokens are sent between GPUs to reach their experts in every layer. That makes inter-GPU communication, typically through a collective library such as NCCL, part of the cost per token. If the weights fit on one card, you avoid that traffic entirely.
Not sure which card fits your model? The GPU recommender can narrow it down, and Aquanode rents GPUs by the hour.
Building on GPUs? Aquanode runs the workload.
Deploy on H100, H200, B200, A100 and MI300X across a multi-provider marketplace, without racking your own hardware or committing to one cloud's spec sheet.
See also
VRAM
VRAM is the memory attached to a GPU that holds the data it works on, and it caps which AI models fit. VRAM vs RAM, how to check yours, and how much AI needs.
Quantization
Quantization stores a model's weights, and sometimes activations, in lower-precision formats like INT8 or 4-bit, cutting VRAM use for a small accuracy cost.
KV Cache
A KV cache stores the key and value tensors of past tokens so an LLM never recomputes them. It grows with context length and batch size, and it eats VRAM.
NCCL (NVIDIA Collective Communications Library)
NCCL is NVIDIA's library for all-reduce and other multi-GPU communication. PyTorch uses it to sync GPUs, and NVLink vs PCIe decides how fast it runs.
HBM (High Bandwidth Memory)
HBM is stacked DRAM packaged beside a GPU die, giving data center cards several TB/s of memory bandwidth. Why LLM inference depends on it, and HBM vs GDDR.
RDMA (Remote Direct Memory Access)
RDMA lets one machine read or write another's memory directly through the network adapter, skipping the remote CPU. It is what multi-node GPU training runs on.