What is NCCL?

Abbreviated NCCL

NCCL (pronounced "nickel"), the NVIDIA Collective Communications Library, is NVIDIA's library for moving and combining data across multiple GPUs with operations such as all-reduce, all-gather, broadcast and reduce-scatter. It works inside one server over NVLink or PCIe and across servers over InfiniBand, RoCE or plain sockets, and it is what PyTorch and most other training frameworks call whenever GPUs need to synchronize.

What a collective does

A collective is an operation that every GPU in a group takes part in at once.

  • All-reduce: every GPU contributes a tensor, and every GPU ends up with the element-wise sum (or another reduction) of all of them. This is the gradient sync in data-parallel training.
  • All-gather: each GPU holds one shard, and every GPU ends up with all the shards. Sharded training uses it to rebuild full weights.
  • Reduce-scatter: the tensors are summed, and each GPU keeps only one shard of the result.
  • Broadcast: one GPU sends the same tensor to all the others, for example to share initial weights.

NCCL detects the machine's topology at startup: which GPUs are joined by NVLink, which sit behind the same PCIe switch, and which network adapter is closest to each GPU. It then picks a route and an algorithm, such as ring or tree, to match. The data movement runs as CUDA kernels on the GPUs themselves, so it can overlap with compute on other streams. Over a network it uses RDMA when the hardware supports it.

How frameworks use it

PyTorch's torch.distributed package has an NCCL backend, which is the one to use for training on NVIDIA GPUs. You select it when you start the process group, usually under torchrun:

import torch.distributed as dist

dist.init_process_group(backend="nccl")

DistributedDataParallel calls all-reduce through it, and sharded training such as FSDP uses all-gather and reduce-scatter. DeepSpeed and Megatron-LM also build on NCCL.

To see what NCCL chose, set NCCL_DEBUG=INFO before launching. It logs the topology and the transport picked for each link at startup, including whether it found InfiniBand or fell back to sockets. NVIDIA also publishes a separate benchmark suite, nccl-tests, whose all_reduce_perf program measures achieved bandwidth.

AMD's counterpart is RCCL, which mirrors NCCL's API for ROCm GPUs; see ROCm vs CUDA for the wider software picture.

What it means when you pick a GPU

An all-reduce is only as fast as the slowest link it has to cross, and inside a node that link is NVLink or PCIe. The H100 SXM has NVLink 4 at 900 GB/s bidirectional, and the A100 SXM has NVLink 3 at 600 GB/s. The L40S and the RTX 4090 have no NVLink at all, so their GPUs can only talk to each other over PCIe. A PCIe Gen 4 x16 slot is about 64 GB/s bidirectional, and Gen 5 about 128 GB/s.

Worked example, at theoretical peak with no overlap with compute. A 7B model in BF16 has 7 billion x 2 bytes = 14 GB of gradients. A ring all-reduce across 8 GPUs sends and receives 2 x (8 - 1) / 8 x 14 = 24.5 GB per GPU. Half of a bidirectional rate is the rate in each direction:

  • NVLink at 900 GB/s: 450 GB/s each way, so 24.5 / 450 = about 0.05 s.
  • PCIe Gen 4 at 64 GB/s: 32 GB/s each way, so 24.5 / 32 = about 0.77 s.

That is roughly 14 times slower on paper, and real PCIe machines route traffic through shared switches, so they often do worse. For multi-GPU training, and for tensor-parallel serving that all-reduces inside every layer, prefer NVLink-connected cards such as the H100. Use PCIe cards for independent jobs, one model per GPU, rather than one job split across them.

Check before you commit. Run nvidia-smi with topo -m on the node: entries starting with NV between GPUs mean NVLink, while PIX, PHB or SYS mean the path goes over PCIe or across CPU sockets. Then run all_reduce_perf from nccl-tests and compare the bandwidth it reports with the link you were promised. For jobs that span nodes, see NVLink vs InfiniBand, and for the on-node side, NVLink vs PCIe.

Building on GPUs? Aquanode runs the workload.

Deploy on H100, H200, B200, A100 and MI300X across a multi-provider marketplace, without racking your own hardware or committing to one cloud's spec sheet.

See also

Submit the job. Everything after that is ours.

Sign up in 60 seconds. Pay for the GPU minutes you actually use.

© 2026 Aquanode. All rights reserved.

All trademarks, logos and brand names are the property of their respective owners.