What is Unified Memory?

Unified memory is a design in which the CPU and the GPU share one pool of memory instead of each having its own. The processor that needs a piece of data reads it where it already is, with no copy across a bus such as PCIe. For running AI models it matters because the model's weights can use the whole pool, not just the slice of memory attached to the GPU.

The phrase covers two different things, and search results mix them up. One is hardware: chips that physically share memory. The other is a software feature of CUDA with the same name. Both are explained below.

Unified memory in hardware

  • Apple M-series. The CPU, GPU and Neural Engine sit on one chip and use a single pool of what Apple calls unified memory. Apple's announcement of the M4 Max lists up to 128GB of unified memory and up to 546GB/s of memory bandwidth.
  • NVIDIA Grace Hopper (GH200). A Grace CPU and a Hopper GPU are joined by NVLink-C2C, which NVIDIA describes as a 900 GB/s coherent link and a "CPU+GPU coherent memory model", 7 times faster than PCIe Gen 5. The GPU keeps its own HBM and can also reach the CPU's LPDDR5X over that link. NVIDIA quotes up to 624GB of combined CPU and GPU fast memory.
  • AMD Instinct MI300A. AMD calls it an APU: CPU cores and GPU compute units on one package. Its data sheet describes 128GB of HBM3 shared coherently between the CPUs and GPUs in a single address space, at 5.3 TB/s peak.

The three are not equal. Apple's pool is one set of memory chips that everything reads at the same speed. Grace Hopper is two pools of different speeds (fast HBM next to the GPU, slower LPDDR5X next to the CPU) that software can address as one. MI300A has the CPU and GPU share the same HBM.

CUDA Unified Memory

In CUDA, unified memory (also called managed memory) is a programming model. You allocate a pointer once and both CPU code and GPU kernels can use it. On a discrete GPU the driver moves the data underneath: pages move to the GPU on demand and can be evicted back to the CPU, which allows a program to use more memory than the card physically has. NVIDIA's programming guide describes hardware-coherent systems such as Grace Hopper as having a combined page table for CPU and GPU, while other systems keep separate ones and rely on page faults to move data. Moving pages is slower than staying in VRAM, so oversubscribing GPU memory this way costs performance.

The QLoRA paper uses this feature for its paged optimizers, which move optimizer state to CPU memory when the GPU runs out and back when it is needed. See QLoRA.

Worked example: a 70B model

Llama 3.1 70B has 70.55 billion parameters (from its published config). At 4 bits per weight, the weights take about 35.3 GB; at 8 bits, 70.6 GB. A rough ceiling on decoding speed at batch size 1 is memory bandwidth divided by the bytes read per token, because every weight is read once per generated token. These are theoretical upper bounds, not benchmarks:

Where the weights sitWeightsBandwidthCeiling
Apple M4 Max, 4-bit, 128GB pool35.3 GB546 GB/sabout 15 tokens/s
H200, 4-bit35.3 GB4,800 GB/sabout 136 tokens/s
H200, FP8 in HBM3e70.6 GB4,800 GB/sabout 68 tokens/s
GH200 with weights read from CPU memory, FP870.6 GB900 at mostabout 13 tokens/s at best

The H200 figures use its 141GB of HBM3e and 4.8 TB/s from the specs on its GPU page. Two things stand out. A unified pool lets a model run that would not fit in one GPU's memory at all, which is a real capability. But what it gives you is capacity, and the speed comes from whichever memory the weights actually live in. A single GPU with HBM is several times faster once the model fits in it.

What it means when you pick a GPU

  • If the model fits in VRAM, keep it there. Weights in HBM read at multiple TB/s. Anything that spills to CPU memory is limited by the link in between: at most 900 GB/s on Grace Hopper (NVIDIA's figure for the whole link, so one direction can be lower) and about 64 GB/s on a PCIe Gen 4 x16 slot (see NVLink vs PCIe). A spill is several times slower on Grace Hopper and well over 50 times slower over PCIe.
  • Use unified memory for capacity, not speed. It suits large models at small batch sizes, long contexts whose KV cache will not fit, and experimentation on a desk. For production throughput, size VRAM for the weights plus cache and pick a card with enough HBM.
  • Shrink the model before you grow the pool. Quantization from 16-bit to 4-bit cuts weights by 4 times and often turns a two-GPU model into a one-GPU one.
  • Check what a card actually has. The GH200 page lists the Superchip's memory and notes it is sold as a module or server, not as a single rentable GPU. Use the VRAM calculator to see whether your model fits in a card's memory, and see pricing for the GPUs Aquanode rents by the hour.

Building on GPUs? Aquanode runs the workload.

Deploy on H100, H200, B200, A100 and MI300X across a multi-provider marketplace, without racking your own hardware or committing to one cloud's spec sheet.

See also

Submit the job. Everything after that is ours.

Sign up in 60 seconds. Pay for the GPU minutes you actually use.

© 2026 Aquanode. All rights reserved.

All trademarks, logos and brand names are the property of their respective owners.