InfiniBand vs Ethernet for GPU Clusters (2026)

Back
Team Aquanode

Team Aquanode

Sarthak Vaish

Updated OCTOBER 8, 2026Published OCTOBER 8, 2026

InfiniBand and Ethernet are the two network fabrics that connect GPU servers in AI clusters. InfiniBand is purpose-built for low-latency RDMA and is NVIDIA's long-standing choice for training clusters. Ethernet, with RDMA over Converged Ethernet and newer AI-tuned stacks such as NVIDIA Spectrum-X and the Ultra Ethernet specification, has closed much of the gap, as of October 2026.

This guide covers:

  • What each fabric is and where it sits relative to NVLink
  • The latest products and published figures: Quantum-X800 InfiniBand, Spectrum-X Ethernet, Ultra Ethernet 1.0
  • Which one fits training, inference and small rented clusters
  • What to check when you rent multi-node GPU capacity

TL;DR

  • Both fabrics carry RDMA traffic between servers. They differ in how the network handles congestion, routing and failures.
  • NVIDIA's current InfiniBand platform, Quantum-X800, has 144 ports of 800 Gb/s per switch. Its Ethernet platform, Spectrum-X, claims a 1.6x gain for AI network performance over off-the-shelf Ethernet. That is NVIDIA's own comparison and the baseline is not fully defined.
  • Ultra Ethernet Consortium published its 1.0 specification on June 11, 2025, an open Ethernet stack for AI and HPC. It is a standard, not a single vendor's product.
  • Verdict: for a few nodes, either works and the choice is mostly made for you by the provider. For large training clusters, InfiniBand is the proven default and AI-tuned Ethernet is the serious alternative, especially where you want multi-vendor gear.

Where the network sits

A GPU cluster has three layers of connectivity. HBM connects a GPU to its own memory. NVLink connects GPUs inside a server or rack. The network connects servers to each other. This post is about the third layer. The comparison between the second and third is in our glossary entry on NVLink vs InfiniBand.

On modern NVIDIA servers, each GPU typically has its own NIC. NVIDIA's DGX B300 page lists eight single-port ConnectX-8 adapters at up to 800 Gb/s each, which can run as InfiniBand or Ethernet. That is the key point: the same adapter family speaks both, so the fabric choice is a decision about switches and software behavior, not about the server.

What RDMA gives both fabrics

Both fabrics exist to carry RDMA traffic: one machine writes directly into another machine's memory without the remote CPU in the path. Libraries such as NCCL use RDMA to run all-reduce, all-gather and similar collective operations across nodes. If the network drops packets or congests unevenly, one slow flow stalls the whole collective, because every GPU waits for the slowest. So the real question about a fabric is how well it behaves under the burst traffic of collective operations, not its headline port speed.

InfiniBand

InfiniBand is a networking standard from the InfiniBand Trade Association, with RDMA as a native feature. NVIDIA is the main vendor of InfiniBand switches and adapters today.

Published figures for the current platform:

  • Quantum-X800. NVIDIA's announcement called Quantum-X800 and Spectrum-X800 the first networking platforms capable of end-to-end 800 Gb/s throughput, using the Quantum Q3400 switch with the ConnectX-8 SuperNIC. NVIDIA's product documentation describes 144 ports of 800 Gb/s per switch with 200 Gb/s-per-lane SerDes.
  • In-network compute. The announcement cites SHARPv4 at 14.4 Tflops of in-network computing, a 9x increase over the previous generation, and 5x higher bandwidth capacity. These are NVIDIA's claims against its own previous generation. SHARP lets the switches do part of a reduction as data flows through, which cuts the data the GPUs have to exchange.
  • Self-healing. NVIDIA's InfiniBand page claims network recovery "5,000X faster than any other software-based solution." Treat that as a vendor claim.
  • Co-packaged optics. NVIDIA offers a liquid-cooled CPO variant, the Q3450-LD, also with 144 ports of 800 Gb/s.

Strengths: mature adaptive routing and congestion control, in-network reduction, and a long record in the largest training clusters. Weaknesses: effectively a single-vendor ecosystem, and a separate fabric from the Ethernet your storage and front-end network probably already use.

Ethernet: RoCE, Spectrum-X and Ultra Ethernet

Plain Ethernet was designed to tolerate loss and rely on higher layers to recover. That is poor for RDMA collectives, which want a lossless, evenly balanced network. Three developments have changed this.

RoCE. RDMA over Converged Ethernet runs RDMA on Ethernet with priority flow control and congestion notification. It works, and it is widely deployed, but it needs careful tuning.

NVIDIA Spectrum-X. An end-to-end Ethernet platform: Spectrum switches plus SuperNICs, with adaptive routing and congestion control. NVIDIA's Spectrum-X page says it accelerates AI network performance by 1.6x over off-the-shelf Ethernet. That page does not define the baseline beyond "OTS Ethernet," and we did not find an independent benchmark confirming the number. Other published figures from the same page:

  • Spectrum-6 switch ASIC for Vera Rubin: 102.4 Tb/s per switch chip with 200G SerDes.
  • ConnectX-8 SuperNIC (Blackwell): 800 Gb/s total throughput over 2x400G, PCIe Gen6.
  • ConnectX-9 SuperNIC (Vera Rubin NVL72): 1,600 Gb/s per GPU over 4x200G SerDes.
  • Multiplane scaling: up to 128K GPUs in two tiers, which NVIDIA says is 64x more than single-plane networks.
  • Spectrum-XGS: NVIDIA claims 1.9x higher NCCL performance in cross-data-center environments, with sites separated by hundreds of kilometers.
  • Customer example on the page: a 100,000-GPU Hopper system.

Ultra Ethernet. The Ultra Ethernet Consortium, a Linux Foundation project founded in mid-2023 by AMD, Arista, Broadcom, Cisco, Eviden, HPE, Intel, Meta and Microsoft, published UEC Specification 1.0 on June 11, 2025. It defines a full Ethernet-based stack for AI and HPC: transport, congestion control, RDMA, link and PHY, and security. It is an open, multi-vendor answer to the single-vendor concern about InfiniBand. We did not find a published performance comparison against InfiniBand, so we do not quote one.

InfiniBand vs Ethernet at a glance

InfiniBand (Quantum-X800)Ethernet (Spectrum-X)Ethernet (Ultra Ethernet)
Standard ownerInfiniBand Trade AssociationIEEE Ethernet, NVIDIA platform on topUltra Ethernet Consortium
Top published port speed800 Gb/s800 Gb/s (ConnectX-8), 1,600 Gb/s per GPU (ConnectX-9)Spec defines the stack; speed follows Ethernet PHYs
RDMANativeRoCE with AI-tuned congestion controlUEC transport
VendorsPrimarily NVIDIANVIDIAMulti-vendor
Headline vendor claim5,000x faster self-healing (NVIDIA's claim)1.6x vs off-the-shelf Ethernet (NVIDIA's claim)No performance claim we could source

Read this table as a map of what each side publishes, not a scoreboard. The claims use different baselines and are not comparable to each other.

Which should you use?

  • Single node (8 GPUs). Neither. NVLink inside the box does the work. Nothing on this page matters until you cross servers.
  • Two to a few nodes, fine-tuning or inference. Either works. Check that the provider gives you RDMA-capable NICs and not plain TCP, and that NCCL is using them. The difference between an RDMA path and a TCP fallback is far larger than the difference between InfiniBand and good Ethernet.
  • Large pretraining runs. InfiniBand is the established default, and its in-network reduction helps collectives. If you are buying your own, Spectrum-X and Ultra Ethernet-based gear are credible options and mean you do not depend on one vendor.
  • Disaggregated or multi-tenant inference. Ethernet is the more natural fit because the same fabric can carry storage and front-end traffic,.
  • Cross-site training. Only one product on this page is aimed at it, Spectrum-XGS, and the 1.9x NCCL claim is NVIDIA's own.

What to check when renting a multi-node cluster

  1. Fabric type and per-GPU bandwidth. Ask for the NIC model and speed per GPU, not just "InfiniBand" or "fast networking."
  2. Is RDMA actually on? Run an NCCL all-reduce test between two nodes and confirm it reports the NIC path, not sockets.
  3. Topology. A non-blocking fat tree behaves very differently from an oversubscribed one.
  4. What the network is not for. Do not size tensor parallelism across nodes. Keep it inside the NVLink domain.

Cost

Network choice shows up in cost indirectly, through idle GPU time during collectives. We do not publish a throughput figure for it, because it depends on the model, the parallelism layout and the software. Measure your own job's step time on both layouts before committing, then multiply by the live hourly price below.

Rent today

Multi-node training and fine-tuning start from these GPUs. Live on-demand prices:

What's next

Spectrum-6, per NVIDIA's Spectrum-X page, doubles per-lane speed to 200G SerDes and reaches 102.4 Tb/s per chip, with the ConnectX-9 SuperNIC at 1,600 Gb/s per GPU for Vera Rubin NVL72. NVIDIA also lists co-packaged optics for Spectrum-X, with claims of 5x better power efficiency and 10x greater mean time between interruptions versus pluggable transceivers. Those are vendor claims.

For every datacenter chip on one page, with status, memory and links to each guide, see the datacenter GPU guide.

FAQ

Is InfiniBand faster than Ethernet?

Top port speeds are the same class: 800 Gb/s on NVIDIA's latest InfiniBand and Ethernet products. The difference is in behavior under collective traffic: congestion control, routing and in-network reduction. Vendor comparisons use different baselines, so there is no single honest "faster" number.

What is Spectrum-X?

NVIDIA's Ethernet platform for AI: Spectrum switches and SuperNICs with adaptive routing and congestion control. NVIDIA claims 1.6x better AI network performance than off-the-shelf Ethernet.

What is Ultra Ethernet?

An open specification from the Ultra Ethernet Consortium, version 1.0 published June 11, 2025, for running AI and HPC traffic over Ethernet, from transport to PHY and security.

Do I need InfiniBand for fine-tuning?

Usually not. Small multi-node fine-tuning is dominated by compute, and what you need is working RDMA, not a particular fabric brand.

Does the network replace NVLink?

No. NVLink connects GPUs inside a server or rack at far higher bandwidth. The network connects servers. See what is NVLink.

Sources

#datacenter gpu#gpu interconnect#infiniband#spectrum-x#ethernet#gpu cluster networking

Submit the job. Everything after that is ours.

Sign up in 60 seconds. Pay for the GPU minutes you actually use.

© 2026 Aquanode. All rights reserved.

All trademarks, logos and brand names are the property of their respective owners.