What is Pipeline Parallelism?

Pipeline parallelism splits a model by depth: GPU 1 holds the first group of layers, GPU 2 the next group, and so on. A batch is cut into small micro-batches that flow through the stages like items on an assembly line, so all the GPUs are busy at once. Only the activations at the boundary between two stages cross the network, which makes it the least bandwidth-hungry way to split one model across GPUs, at the price of some idle time.

It is also called inter-layer model parallelism. The scheme in use today descends from GPipe (Huang et al., 2018).

How it works

Without micro-batches, only one stage would be working at any moment while the others wait. GPipe divides each mini-batch into M micro-batches and sends them through the K stages one after another. In the forward pass, stage 2 starts on micro-batch 1 as soon as stage 1 hands it over, while stage 1 moves on to micro-batch 2. The backward pass runs the same way in reverse, and gradients from all micro-batches are summed before one weight update, so the result matches ordinary training on the whole mini-batch.

The unavoidable idle time is the bubble: the start of the pass, when later stages have nothing yet, and the end, when earlier stages have finished. GPipe gives the bubble overhead as proportional to (K - 1) / (M + K - 1). For stages of equal length, that is also the fraction of time each GPU sits idle. It shrinks as the number of micro-batches grows relative to the number of stages, and GPipe reports the overhead was negligible in its experiments when M is at least 4 x K.

GPipe also re-computes activations during the backward pass instead of storing them, so each stage keeps only the activations at its boundary. That trades extra compute for memory.

Worked example: Llama 3.1 70B on 4 GPUs

From the model's published config.json: 80 layers, hidden size 8192, feed-forward size 28,672, 70.55 billion parameters, 141.1 GB in BF16.

  • Memory. With K = 4 stages of 20 layers each, a stage holds about 17.1 billion parameters of layer weights, 34.2 GB in BF16, plus the embedding on the first stage and the output head on the last, which are about 1.05 billion parameters (2.1 GB) each. A stage therefore needs roughly 35 to 37 GB for weights.
  • Bubble. Using (K - 1) / (M + K - 1) with K = 4:
Micro-batches (M)Idle fraction
175.0%
442.9%
1615.8%
328.6%

These come from the formula for equal-length stages and ignore the effect of scheduling tricks, so treat them as estimates. With one micro-batch the pipeline is no better than a single GPU working while three wait.

  • Traffic. The only thing sent between stages is the activation tensor. For a micro-batch of 4,096 tokens that is 4,096 x 8,192 x 2 bytes = 67.1 MB per boundary, and a pass crosses 3 boundaries. Even over a PCIe Gen 4 link at 32 GB/s per direction, one hop takes about 2 ms. Compare tensor parallelism, which does 160 all-reduces of the same size in one forward pass of this model.

That bandwidth gap is the reason to choose pipeline parallelism when the GPUs are not on NVLink. GPipe's authors make the same point: since only boundary activations move, it scales efficiently even on accelerators without high-speed interconnects.

Versus the other parallelism types

  • Data parallelism replicates the full model and needs it to fit on one GPU. Pipeline parallelism removes that requirement and the two combine: run several pipeline replicas on different slices of the data.
  • Tensor parallelism splits within each layer and communicates inside every layer. It cuts latency per token, while a pipeline raises throughput but not single-request latency, since a token still passes through every stage in turn.
  • Pipelines can combine with tensor parallelism and sharding too: the Megatron-LM paper proposes combining intra-layer and inter-layer parallelism for models too big for one server.

What it means when you pick a GPU

Pipeline parallelism is the tool for splitting a model across GPUs that are not tightly connected: PCIe-only cards, or separate servers linked by Ethernet or InfiniBand. Stage-to-stage traffic is small, so a link that would cripple tensor parallelism is often fine here.

What you pay for is the bubble, so it suits large batches of work. Training can feed many micro-batches into the pipe, whereas interactive serving with a few concurrent requests cannot, which leaves the stages idle. Balance the stages: they run at the speed of the slowest, so give the stage with the embedding or output head fewer layers. And size memory per stage, not for the whole model: use the VRAM calculator for one stage's share of the weights plus its activations.

For a model that needs more than one 8-GPU server, one option is tensor parallelism inside each server and pipeline stages between servers, the hybrid the Megatron-LM paper points to. See GPU clusters and our guide to AI compute clusters, and pricing for what Aquanode rents by the hour.

Building on GPUs? Aquanode runs the workload.

Deploy on H100, H200, B200, A100 and MI300X across a multi-provider marketplace, without racking your own hardware or committing to one cloud's spec sheet.

See also

Submit the job. Everything after that is ours.

Sign up in 60 seconds. Pay for the GPU minutes you actually use.

© 2026 Aquanode. All rights reserved.

All trademarks, logos and brand names are the property of their respective owners.