What is Pipeline Parallelism?
Pipeline parallelism splits a model by depth: GPU 1 holds the first group of layers, GPU 2 the next group, and so on. A batch is cut into small micro-batches that flow through the stages like items on an assembly line, so all the GPUs are busy at once. Only the activations at the boundary between two stages cross the network, which makes it the least bandwidth-hungry way to split one model across GPUs, at the price of some idle time.
It is also called inter-layer model parallelism. The scheme in use today descends from GPipe (Huang et al., 2018).
How it works
Without micro-batches, only one stage would be working at any moment while the others wait. GPipe divides each mini-batch into M micro-batches and sends them through the K stages one after another. In the forward pass, stage 2 starts on micro-batch 1 as soon as stage 1 hands it over, while stage 1 moves on to micro-batch 2. The backward pass runs the same way in reverse, and gradients from all micro-batches are summed before one weight update, so the result matches ordinary training on the whole mini-batch.
The unavoidable idle time is the bubble: the start of the pass, when later stages have nothing yet, and the end, when earlier stages have finished. GPipe gives the bubble overhead as proportional to (K - 1) / (M + K - 1). For stages of equal length, that is also the fraction of time each GPU sits idle. It shrinks as the number of micro-batches grows relative to the number of stages, and GPipe reports the overhead was negligible in its experiments when M is at least 4 x K.
GPipe also re-computes activations during the backward pass instead of storing them, so each stage keeps only the activations at its boundary. That trades extra compute for memory.
Worked example: Llama 3.1 70B on 4 GPUs
From the model's published config.json: 80 layers, hidden size 8192, feed-forward size 28,672, 70.55 billion parameters, 141.1 GB in BF16.
- Memory. With K = 4 stages of 20 layers each, a stage holds about 17.1 billion parameters of layer weights, 34.2 GB in BF16, plus the embedding on the first stage and the output head on the last, which are about 1.05 billion parameters (2.1 GB) each. A stage therefore needs roughly 35 to 37 GB for weights.
- Bubble. Using (K - 1) / (M + K - 1) with K = 4:
| Micro-batches (M) | Idle fraction |
|---|---|
| 1 | 75.0% |
| 4 | 42.9% |
| 16 | 15.8% |
| 32 | 8.6% |
These come from the formula for equal-length stages and ignore the effect of scheduling tricks, so treat them as estimates. With one micro-batch the pipeline is no better than a single GPU working while three wait.
- Traffic. The only thing sent between stages is the activation tensor. For a micro-batch of 4,096 tokens that is 4,096 x 8,192 x 2 bytes = 67.1 MB per boundary, and a pass crosses 3 boundaries. Even over a PCIe Gen 4 link at 32 GB/s per direction, one hop takes about 2 ms. Compare tensor parallelism, which does 160 all-reduces of the same size in one forward pass of this model.
That bandwidth gap is the reason to choose pipeline parallelism when the GPUs are not on NVLink. GPipe's authors make the same point: since only boundary activations move, it scales efficiently even on accelerators without high-speed interconnects.
Versus the other parallelism types
- Data parallelism replicates the full model and needs it to fit on one GPU. Pipeline parallelism removes that requirement and the two combine: run several pipeline replicas on different slices of the data.
- Tensor parallelism splits within each layer and communicates inside every layer. It cuts latency per token, while a pipeline raises throughput but not single-request latency, since a token still passes through every stage in turn.
- Pipelines can combine with tensor parallelism and sharding too: the Megatron-LM paper proposes combining intra-layer and inter-layer parallelism for models too big for one server.
What it means when you pick a GPU
Pipeline parallelism is the tool for splitting a model across GPUs that are not tightly connected: PCIe-only cards, or separate servers linked by Ethernet or InfiniBand. Stage-to-stage traffic is small, so a link that would cripple tensor parallelism is often fine here.
What you pay for is the bubble, so it suits large batches of work. Training can feed many micro-batches into the pipe, whereas interactive serving with a few concurrent requests cannot, which leaves the stages idle. Balance the stages: they run at the speed of the slowest, so give the stage with the embedding or output head fewer layers. And size memory per stage, not for the whole model: use the VRAM calculator for one stage's share of the weights plus its activations.
For a model that needs more than one 8-GPU server, one option is tensor parallelism inside each server and pipeline stages between servers, the hybrid the Megatron-LM paper points to. See GPU clusters and our guide to AI compute clusters, and pricing for what Aquanode rents by the hour.
Building on GPUs? Aquanode runs the workload.
Deploy on H100, H200, B200, A100 and MI300X across a multi-provider marketplace, without racking your own hardware or committing to one cloud's spec sheet.
See also
Data Parallelism
Data parallelism copies the whole model onto every GPU, splits each batch between them and averages gradients with an all-reduce. Simple, but memory-hungry.
Tensor Parallelism
Tensor parallelism splits each layer's matrix multiplications across several GPUs that work on every token together. It needs NVLink-class bandwidth.
NCCL (NVIDIA Collective Communications Library)
NCCL is NVIDIA's library for all-reduce and other multi-GPU communication. PyTorch uses it to sync GPUs, and NVLink vs PCIe decides how fast it runs.
NVLink vs InfiniBand
NVLink connects GPUs inside one server; InfiniBand connects servers to each other. Where the two interconnects overlap, where they don't, and why big clusters run both.
RDMA (Remote Direct Memory Access)
RDMA lets one machine read or write another's memory directly through the network adapter, skipping the remote CPU. It is what multi-node GPU training runs on.
VRAM
VRAM is the memory attached to a GPU that holds the data it works on, and it caps which AI models fit. VRAM vs RAM, how to check yours, and how much AI needs.