What Is NVLink? NVLink 5, NVSwitch and NVL72 (2026)

Back
Team Aquanode

Team Aquanode

Sarthak Vaish

Updated OCTOBER 8, 2026Published OCTOBER 8, 2026

NVLink is NVIDIA's proprietary high-bandwidth link that connects GPUs to each other directly, so a group of GPUs can behave like one large accelerator. On Blackwell, NVIDIA says each GPU has 18 NVLink connections at 100 GB/s each, for 1.8 TB/s per GPU, and NVSwitch chips extend that to 72 GPUs in one rack, as of October 2026.

This guide covers:

  • What NVLink is and why training and inference on large models depend on it
  • Bandwidth by generation, from Hopper (NVLink 4) to Vera Rubin (NVLink 6)
  • What NVSwitch and the NVLink Switch do, and what NVL72 means
  • How NVLink differs from PCIe and from InfiniBand
  • When NVLink matters for your workload and when it does not

TL;DR

  • NVLink is a GPU-to-GPU fabric inside a server or a rack. It is not a network between servers. Servers talk over InfiniBand or Ethernet, which is a separate layer.
  • Per-GPU NVLink bandwidth, per NVIDIA's NVLink page: 900 GB/s on Hopper (NVLink 4), 1,800 GB/s on Blackwell (NVLink 5), and 3,000 GB/s on Vera Rubin (NVLink 6, listed as preliminary).
  • NVSwitch is the switch chip that makes any-to-any NVLink possible. An 8-GPU HGX board uses it inside one server. The NVLink Switch in an NVL72 rack links 72 Blackwell GPUs, with 130 TB/s of total NVLink bandwidth in a GB300 NVL72, per NVIDIA.
  • Verdict: if your job shards one model across several GPUs, NVLink bandwidth is one of the first numbers to check. If each GPU serves a whole model on its own, it barely matters.

What NVLink actually is

A GPU's own memory is fast. A B200 or H100 reads its high bandwidth memory at several terabytes per second. The problem starts the moment a model no longer fits on one GPU. Layers, attention heads or experts get split across GPUs, and every step they have to exchange activations or gradients. If that traffic crosses the PCIe bus, it is limited to a small fraction of what the GPU can read locally, and the GPUs sit idle waiting on each other.

NVLink removes that bottleneck. It is a set of point-to-point links, run directly between GPUs, that lets one GPU read and write another GPU's memory without going through the CPU or the PCIe slot. Our glossary has the short version in NVLink vs PCIe. This post is the longer one: generations, switches, racks and where the limits are.

One useful way to hold it in your head: HBM is the bandwidth between a GPU and its own memory, NVLink is the bandwidth between a GPU and its neighbors, and the network (InfiniBand or Ethernet) is the bandwidth between servers. Each step out is slower and larger in scale.

NVLink bandwidth by generation

NVIDIA's NVLink page lists the following per-GPU figures and GPU domain sizes. The page marks the Vera Rubin row as preliminary and subject to change.

GenerationGPU architectureBandwidth per GPULinks per GPULargest NVLink domain listedTotal NVLink bandwidth of that domain
NVLink 4Hopper (H100, H200)900 GB/s188 GPUs7.2 TB/s
NVLink 5Blackwell (B200, B300)1,800 GB/s1872 GPUs (NVL72)130 TB/s
NVLink 6Vera Rubin3,000 GB/s3672 GPUs (NVL72)216 TB/s

Two things stand out. First, Blackwell kept 18 links and doubled the speed of each one, to 100 GB/s, which is how NVIDIA gets from 900 GB/s to 1.8 TB/s. Vera Rubin doubles the link count to 36 instead. Second, the domain size jumped. On Hopper the NVLink domain is the 8 GPUs on one board. On Blackwell and Rubin, NVIDIA's rack-scale systems extend it to 72 GPUs.

For scale: 1.8 TB/s per Blackwell GPU is about 14 times the 128 GB/s of PCIe Gen5 (1,800 divided by 128, our arithmetic from two NVIDIA pages). NVIDIA's H200 page lists PCIe Gen5 at 128 GB/s. These are NVIDIA's numbers for peak link bandwidth, not measured application throughput.

NVSwitch and the NVLink Switch

Links alone give you point-to-point connections. With 8 GPUs and 18 links each, you could wire pairs together, but you could not give every GPU full bandwidth to every other GPU at once. NVSwitch solves this. It is a dedicated switch chip that every GPU connects to, so any GPU can talk to any other at full link speed. We cover the chip in the glossary entry on NVSwitch.

There are two physical forms of this:

  • Inside a server (HGX and DGX). An HGX B200 or HGX B300 board holds 8 GPUs and NVSwitch chips. NVIDIA's HGX page lists 1.8 TB/s of NVLink bandwidth per GPU and 14.4 TB/s total for both boards, which is 8 GPUs times 1.8 TB/s. The DGX B300 page also lists 14.4 TB/s of aggregate NVLink bandwidth, through NVLink Switch hardware in the box.
  • Across a rack (NVL72). In rack-scale systems such as the GB200 NVL72 and GB300 NVL72, NVLink Switch trays connect 72 GPUs in a non-blocking fabric. NVIDIA's GB300 NVL72 page lists 130 TB/s of NVLink bandwidth for the rack, and 37 TB of fast memory (20 TB of GPU memory plus 17 TB of CPU LPDDR5X). The Vera Rubin NVL72 is listed at 216 TB/s.

NVIDIA's Blackwell launch release says fifth-generation NVLink supports "up to 576 GPUs" in one communication domain. That figure is a fabric capability, and as of this writing the shipping rack product on NVIDIA's page is the 72-GPU rack. NVIDIA's current NVLink page lists 8 and 72 GPU domains, not 576, so treat 576 as a design limit, not a product you can buy. We do not state the switch chip's port count or switching capacity because we could not confirm them on an NVIDIA page.

The NVLink page also says each NVLink Switch has engines for SHARP in-network reductions and multicast. In plain terms, the switch can do part of the math of a collective operation (an all-reduce, for example) while data passes through it, instead of making every GPU do all of it.

NVLink vs PCIe vs InfiniBand

People mix these up because all three show up on a cluster spec sheet. They sit at different layers.

NVLinkPCIeInfiniBand / Ethernet
ConnectsGPU to GPU (and GPU to Grace CPU on superchips)GPU to CPU, NICs, SSDsServer to server
Who makes itNVIDIA, proprietaryOpen standardOpen standards (InfiniBand is NVIDIA's main product; Ethernet is multi-vendor)
Per-GPU figure1.8 TB/s on Blackwell128 GB/s on Gen5 x16, per NVIDIA's H200 page800 Gb/s per port on the latest NVIDIA gear (see below)
ScopeOne server or one rackOne serverA whole data center

The network figure is in bits, the others are in bytes. 800 Gb/s is 100 GB/s. NVIDIA's DGX B300 page lists eight ConnectX-8 NICs at up to 800 Gb/s each for scale-out. So a DGX B300 has 14.4 TB/s of NVLink inside the box and a fraction of that going out to other boxes. This gap is why training software places the chattiest parallelism (tensor parallelism) inside the NVLink domain and the less chatty kinds (data and pipeline parallelism) across the network.

For the cluster-network side of the story, see our glossary entry on NVLink vs InfiniBand and the post on InfiniBand vs Ethernet for GPU clusters.

Why NVL72 changed the shape of inference

Before NVL72, the NVLink domain topped out at 8 GPUs. A model that needed more than 8 GPUs' memory had to cross the slower network, and the network became the limit. Rack-scale NVLink raises that ceiling from 8 to 72 GPUs. That matters most for two workloads:

  • Large mixture-of-experts models. Experts live on different GPUs, and every token is routed to a few of them. That is an all-to-all pattern, which suffers on a slower fabric.
  • Long-context and high-throughput serving. A bigger NVLink domain lets you shard one model's weights and KV cache over more GPUs while keeping the traffic on the fast fabric.

We are not publishing a benchmark here, because the honest numbers depend on the model, the batch size and the software stack. NVIDIA publishes MLPerf results and its own claims for NVL72 systems; read those with the vendor label in mind and check which model and precision they used.

Does NVLink matter for your workload?

A short checklist:

  • Single-GPU inference of a model that fits in one GPU's memory. NVLink does not matter. A single H100 or B200 never uses it.
  • Multi-GPU inference of a big model on one 8-GPU server. NVLink matters a lot. This is the HGX case. Check that you are renting an SXM board with NVSwitch and not PCIe cards. Our glossary entry on NVLink vs PCIe explains the difference between the two form factors.
  • Fine-tuning with FSDP or tensor parallelism across 8 GPUs. NVLink matters. Collective operations such as all-gather and reduce-scatter run through NCCL, which uses NVLink when it is available.
  • Multi-node training. Both layers matter. NVLink carries tensor-parallel traffic inside each node, and the network carries the rest.
  • Embarrassingly parallel work (many independent small jobs). NVLink does not matter.

One practical point: NVLink is also a pricing signal. A PCIe-attached 8-GPU server can look cheaper per GPU, but for a workload that exchanges data every step, the link can cost you more in lost time than the hardware saves.

NVLink Fusion and the open alternative

NVIDIA's NVLink page says NVLink Fusion pairs NVIDIA technology with semi-custom ASICs or CPUs, so hyperscalers can build hybrid infrastructure on NVLink and rack-scale architecture. The alternative approach is an open standard. UALink, backed by AMD, Google, Meta, Microsoft and others, targets the same scale-up role. We compare the two in UALink vs NVLink.

Rent today

NVLink-connected 8-GPU boxes are the standard configuration for large-model work. Here are the live on-demand prices for a Blackwell and a Hopper GPU, both with NVLink in their SXM form.

What's next

NVIDIA's NVLink page lists NVLink 6 for Vera Rubin at 3,000 GB/s per GPU and 216 TB/s per NVL72 rack, marked preliminary. NVIDIA's Rubin technical blog and HGX page differ, giving 3.6 TB/s per GPU and 260 TB/s per rack; we show the NVLink page and product page values in the table. The Vera Rubin NVL72 product page labels the system "Available Now" and says Vera Rubin is ramping into full production. Check our Rubin guide for what that means for rentable capacity.

For every datacenter chip on one page, with status, memory and links to each guide, see the datacenter GPU guide.

FAQ

Is NVLink faster than PCIe?

Yes. NVIDIA lists 1.8 TB/s per Blackwell GPU over NVLink 5, against 128 GB/s for PCIe Gen5, which is the figure on its H200 page. That is about 14 times (our arithmetic).

What is the difference between NVLink and NVSwitch?

NVLink is the link. NVSwitch is the switch chip that joins many NVLinks so every GPU can reach every other GPU at full speed. A server with 8 GPUs and no switch can only connect some pairs directly.

Can I use NVLink between GPUs in different servers?

Not in a normal server. Between servers you use InfiniBand or Ethernet. The exception is a rack-scale system like NVL72, where NVLink Switch trays connect 72 GPUs across the rack's compute trays.

Do AMD and Intel GPUs have NVLink?

No. NVLink is NVIDIA's technology. AMD uses Infinity Fabric inside a node and is moving to UALink-based scale-up, which we cover in UALink vs NVLink.

Does NVLink matter for inference?

It matters when one model is split across GPUs, which is common for large models and mixture-of-experts. For a model that fits in one GPU, the GPUs do not talk to each other during a request.

Sources

#datacenter gpu#gpu interconnect#nvlink#nvswitch#nvlink 5#nvl72

Submit the job. Everything after that is ours.

Sign up in 60 seconds. Pay for the GPU minutes you actually use.

© 2026 Aquanode. All rights reserved.

All trademarks, logos and brand names are the property of their respective owners.