What is Tensor Parallelism?

Tensor parallelism splits the weight matrices inside each layer of a model across several GPUs, so every GPU computes a slice of every layer and the slices are combined with a collective operation. It is how a model too large for one GPU can run with all its GPUs working on the same token at once, which cuts latency as well as memory. The cost is that the GPUs must exchange data several times per layer, so it only works well over very fast links.

It is also called intra-layer model parallelism. The standard recipe comes from NVIDIA's Megatron-LM paper (Shoeybi et al., 2019).

How it works

A transformer layer is mostly two things: attention and a feed-forward block. Megatron-LM splits both so that each GPU needs no communication in the middle:

  • Feed-forward block. The first weight matrix is split by columns, so each GPU produces its own slice of the hidden activations. The second matrix is split by rows, so each GPU consumes exactly the slice it produced. The partial outputs are then summed with one all-reduce.
  • Attention. The attention heads are divided among the GPUs. Each GPU runs its own heads independently and the output projection is split by rows, again ending in one all-reduce.

The paper counts the result as two all-reduces in the forward pass and two in the backward pass per transformer layer. The all-reduce is the same collective described under NCCL, and it sits directly on the critical path: the next layer cannot start until it finishes.

Worked example: Llama 3.1 70B on 8 GPUs

From the model's published config.json: 80 layers, hidden size 8192, feed-forward size 28,672, 64 attention heads and 8 key/value heads, which comes to 70.55 billion parameters.

  • Memory. In BF16 the weights are 141.1 GB. Split 8 ways, each GPU holds 17.6 GB. Each GPU also gets 64 / 8 = 8 query heads, 8 / 8 = 1 key/value head, and 28,672 / 8 = 3,584 feed-forward columns per layer.
  • Traffic. Each all-reduce moves one activation tensor of tokens x hidden size x 2 bytes. For a 4,096-token prompt that is 4,096 x 8,192 x 2 = 67.1 MB. A ring all-reduce over 8 GPUs sends 2 x 7 / 8 x 67.1 = 117.4 MB per GPU. With 2 all-reduces in each of 80 layers, a single forward pass does 160 of them: 18.8 GB per GPU.
  • Time at theoretical peak, with half the bidirectional link rate per direction as in the NCCL example:
Link inside the nodePer directionCommunication for one 4,096-token pass
NVLink 4 (H100), 900 GB/s450 GB/sabout 0.04 s
PCIe Gen 4 x16, 64 GB/s32 GB/sabout 0.59 s

For scale, the matrix math for that pass is about 2 x 70.55 billion x 4,096 = 5.8 x 10^14 FLOPs. Spread over 8 H100s at their 989 TFLOPS dense BF16 peak, that is about 0.07 s of compute. So even over NVLink the communication is comparable to the compute unless it overlaps, and over PCIe it is about 8 times the compute. These are upper-bound arithmetic, not benchmarks, but the ratio is why tensor parallelism is kept inside a node. For single-token decoding the tensors are tiny (8,192 x 2 bytes = 16 KB), so the cost is latency from 160 collectives per token rather than bandwidth.

The Megatron-LM paper reports the same pattern in practice: on servers with 300 GB/s between GPUs through NVSwitch and 100 GB/s between servers, its 8.3-billion-parameter model reached 77% of linear scaling with 8-way model parallelism inside one server.

Versus the other parallelism types

  • Data parallelism copies the whole model and splits the batch. Tensor parallelism splits the model and shares the batch.
  • Pipeline parallelism gives each GPU a block of whole layers and sends activations between them. It needs far less bandwidth than tensor parallelism. The Megatron-LM paper itself suggests combining intra-layer (tensor) and inter-layer (pipeline) parallelism for models too large for one 16-GPU server.
  • FSDP shards parameters but gathers whole layers before computing them, so each GPU still does full-width matrix math. Tensor parallelism never rebuilds a full layer.

What it means when you pick a GPU

Use tensor parallelism when a model's weights do not fit on one GPU, or when you need lower latency per token than one GPU gives, and only across GPUs joined by NVLink. The fabric matters more than the GPU count. An H100 has NVLink 4 at 900 GB/s, an A100 SXM has NVLink 3 at 600 GB/s, and cards such as the L40S and RTX 4090 have no NVLink and fall back to PCIe. Eight GPUs all-to-all through an NVSwitch give each one its full link to any other, while GPUs joined only by pairwise bridges do not.

Choose the degree to match what splits cleanly: the number of attention heads and key/value heads has to divide evenly, so 8 is natural for a model with 8 key/value heads like this one. Check the topology with nvidia-smi topo -m and measure with nccl-tests before a long run (see NVLink vs PCIe). For the multi-GPU node itself, see GPU clusters and our guide to AI compute clusters. Aquanode rents GPUs by the hour.

Building on GPUs? Aquanode runs the workload.

Deploy on H100, H200, B200, A100 and MI300X across a multi-provider marketplace, without racking your own hardware or committing to one cloud's spec sheet.

See also

Submit the job. Everything after that is ours.

Sign up in 60 seconds. Pay for the GPU minutes you actually use.

© 2026 Aquanode. All rights reserved.

All trademarks, logos and brand names are the property of their respective owners.