What is NVSwitch?
NVSwitch is NVIDIA's switch chip for NVLink. It lets every GPU in a group talk to every other GPU at full NVLink bandwidth at the same time, so eight or more GPUs behave like one big accelerator instead of a set of cards joined by a few direct cables. NVIDIA markets it, together with NVLink, as "NVLink Switch".
The problem it solves
NVLink is a point-to-point link: each GPU has a fixed number of NVLink ports, and each port connects to one other device. Wire GPUs directly to each other and the ports have to be divided among the peers. With 8 GPUs, each GPU has 7 peers, so the fixed link budget is split 7 ways. Every pair gets a thin slice, and no single pair can ever use the GPU's full bandwidth.
A switch fixes this. Each GPU connects all its links to the switch chips, and the switches route traffic between any pair. One GPU can send its entire NVLink bandwidth to a single peer, or spread it across all seven, and the fabric does not care.
What NVIDIA publishes
From NVIDIA's NVLink and NVLink Switch specifications page:
| NVLink 4 (Hopper) | NVLink 5 (Blackwell) | |
|---|---|---|
| Bandwidth per GPU | 900 GB/s | 1,800 GB/s |
| Maximum links per GPU | 18 | 18 |
| NVLink domain sizes | 8 | 8 or 72 |
| GPU-to-GPU bandwidth through the switch | 900 GB/s | 1,800 GB/s |
| Total aggregate bandwidth | 7.2 TB/s | 130 TB/s (NVL72 rack) |
The aggregate numbers follow from the per-GPU ones: 8 x 900 GB/s = 7.2 TB/s, and 72 x 1.8 TB/s is about 130 TB/s. The 72-GPU domain is the main change in the Blackwell generation: NVLink 5 switches connect up to 72 GPUs in a single rack-scale NVLink domain, not only the 8 inside a server. NVIDIA's page also lists a sixth generation (NVLink 6) as preliminary specifications.
Worked example: why a switch matters for 8 GPUs
An H100 has 18 NVLink links totalling 900 GB/s, so each link carries 50 GB/s (900 / 18). Compare two ways of connecting 8 of them, using arithmetic from those published figures:
- Direct wiring, no switch. 18 links over 7 peers is 2 links per peer (using 14, with 4 left over), so each pair gets 2 x 50 = 100 GB/s. A GPU talking to one peer reaches only a ninth of its 900 GB/s.
- Through NVSwitch. Each GPU connects all 18 links to the switches, and NVIDIA's table lists 900 GB/s of GPU-to-GPU bandwidth. A GPU can send its entire 900 GB/s to any one peer, or divide it across all seven.
The effect is largest for collectives. In an all-reduce, every GPU talks to others constantly, and under tensor parallelism the same 160 all-reduces per forward pass land on the critical path (see the worked example in tensor parallelism). The Megatron-LM authors ran 8-way tensor parallelism inside DGX-2H servers, whose GPUs talked to each other at 300 GB/s through NVSwitch, versus 100 GB/s between servers over InfiniBand, and they reported 77% of linear scaling for that setup. The jump in bandwidth at the server boundary is also why parallelism is arranged to keep the chattiest traffic inside the NVLink domain. The traffic across servers goes over InfiniBand or RoCE using RDMA.
NVSwitch versus a bridge
Not every NVLink GPU has a switch. Some PCIe cards connect in pairs or small groups with a physical NVLink bridge, which is point-to-point only. The PCIe-form H200 NVL, for example, uses an NVLink bridge for 2 to 4 GPUs, not the full SXM NVLink fabric. That is much better than PCIe between the bridged cards, but they cannot reach non-bridged GPUs over it. A switched fabric is what lets all 8 (or 72) GPUs reach each other at full speed.
What it means when you pick a GPU
If you plan to run one job across many GPUs, tensor-parallel serving, large-model training or anything with frequent all-reduces, the interconnect topology is part of the product you are renting. Look for 8-GPU SXM-class nodes (H100, H200, B200 and A100 SXM have NVLink at 900, 900, 1,800 and 600 GB/s respectively in our GPU specs). Cards with no NVLink, such as the L40S and RTX 4090, only talk over PCIe at about 64 GB/s per slot at Gen 4.
Check what you got rather than trusting the listing. Run nvidia-smi topo -m on the node (see nvidia-smi). A switched node shows the same NVLink entry for every pair of GPUs, while a bridged or PCIe-only node shows NVLink to a partner only, or PIX/PHB/SYS entries. Then run nccl-tests' all_reduce_perf and compare the bandwidth with the figures above (see NCCL). If your job never talks across GPUs, as with one independent model per card, NVSwitch buys you nothing and a PCIe node does the same work. For multi-GPU nodes on Aquanode, see GPU clusters and pricing.
Building on GPUs? Aquanode runs the workload.
Deploy on H100, H200, B200, A100 and MI300X across a multi-provider marketplace, without racking your own hardware or committing to one cloud's spec sheet.
See also
NVLink vs PCIe
NVLink and PCIe both move data in and out of a GPU, but at very different scales. Where each interconnect wins on bandwidth, latency, cost, and compatibility.
NVLink vs InfiniBand
NVLink connects GPUs inside one server; InfiniBand connects servers to each other. Where the two interconnects overlap, where they don't, and why big clusters run both.
NCCL (NVIDIA Collective Communications Library)
NCCL is NVIDIA's library for all-reduce and other multi-GPU communication. PyTorch uses it to sync GPUs, and NVLink vs PCIe decides how fast it runs.
Tensor Parallelism
Tensor parallelism splits each layer's matrix multiplications across several GPUs that work on every token together. It needs NVLink-class bandwidth.
GPU Architecture
How NVIDIA GPUs are actually built, from Graphics Processing Clusters and Streaming Multiprocessors down to the memory hierarchy, and how Ampere, Hopper, and Blackwell differ.
RDMA (Remote Direct Memory Access)
RDMA lets one machine read or write another's memory directly through the network adapter, skipping the remote CPU. It is what multi-node GPU training runs on.