What is RDMA (Remote Direct Memory Access)?

Abbreviated RDMA

RDMA (Remote Direct Memory Access) lets one machine read or write another machine's memory directly through the network adapter, without involving the remote CPU or its operating system kernel. Skipping the CPU and kernel on the data path cuts latency and frees CPU time, which is why RDMA networks connect the servers in multi-node AI training clusters.

With ordinary TCP sockets, every message is copied between application memory and kernel buffers, and the CPU handles protocol work on both ends. With RDMA, the application registers a region of memory with the network adapter once, then posts transfer requests straight to the adapter hardware. The adapters move the bytes between the two registered regions on their own, with no extra copies and no kernel on the fast path.

RDMA transports: InfiniBand and RoCE

RDMA is a capability, not a single network. It runs over two main transports:

  • InfiniBand: a fabric designed around RDMA from the start, with its own switches, cabling and adapters. Its current NDR generation runs at 400 Gb/s per port.
  • RoCE (RDMA over Converged Ethernet): the same RDMA operations carried over Ethernet. RoCE v2 runs over UDP/IP, so it can cross routed networks. RDMA transports generally expect a network that rarely drops packets, so the Ethernet fabric usually has to be configured for that.

InfiniBand and RoCE give the same programming model to software, which is why libraries such as NCCL can use either.

GPUDirect RDMA

Plain RDMA moves data between host memory on two machines. In a GPU server the data you care about sits in GPU memory, so without help it takes an extra hop: GPU memory to a host buffer, then host buffer to the network adapter.

GPUDirect RDMA is NVIDIA's technology that lets the network adapter read and write GPU memory directly over PCIe, removing the host-memory hop. It needs a supported NVIDIA GPU, a supported adapter and matching drivers, and NCCL uses it automatically when the stack allows.

Why this matters for training: in data-parallel training, every GPU computes gradients on its own slice of the batch, and then all GPUs run an all-reduce to average those gradients before the next step. Across several machines, that all-reduce crosses the network on every single step, so network bandwidth and latency set how long each step waits.

What it means when you pick a GPU

A job that fits on one node does not need RDMA, however many GPUs it uses. GPUs inside one node talk to each other over NVLink or PCIe, and nothing crosses a network. Only a job that spans several nodes depends on the fabric between them.

Here is why link speed matters, with the arithmetic. A 7B model in BF16 produces 7 billion x 2 bytes = 14 GB of gradients per step. A bandwidth-optimal ring all-reduce over many participants moves about 2 x 14 = 28 GB through each participant's link. Over a 400 Gb/s port (50 GB/s each way) that is 28 / 50 = 0.56 s per step. Over a 25 Gb/s link (3.125 GB/s) it is 28 / 3.125 = about 9 s per step. Compare either number with how long a step takes to compute on your GPUs.

This is a simplified lower bound: it assumes one adapter per participant and no overlap with compute. Real clusters use several adapters per node, overlap communication with computation, and shard state to cut traffic. The point is the ratio between link speeds, not the exact time.

Before renting for a multi-node job, check these things:

  • The fabric type (InfiniBand or RoCE) and the per-node bandwidth, including how many adapters each node has.
  • Whether GPUDirect RDMA is supported on those nodes.
  • A real all-reduce measurement across two nodes, using the NCCL test tools, before you commit to a long run.

If your workload fits on one node, the GPU recommender can help you pick a card without needing any of this. For how the intra-node and inter-node links divide the work, see NVLink vs InfiniBand.

Building on GPUs? Aquanode runs the workload.

Deploy on H100, H200, B200, A100 and MI300X across a multi-provider marketplace, without racking your own hardware or committing to one cloud's spec sheet.

See also

Submit the job. Everything after that is ours.

Sign up in 60 seconds. Pay for the GPU minutes you actually use.

© 2026 Aquanode. All rights reserved.

All trademarks, logos and brand names are the property of their respective owners.