What is HBM (High Bandwidth Memory)?

Abbreviated HBM

HBM (High Bandwidth Memory) is DRAM built as vertical stacks of memory chips and mounted in the same package as the GPU die, which gives it far more bandwidth than the GDDR memory on consumer cards. Data center GPUs such as the H100, H200 and MI300X use it because LLM inference spends most of its time waiting on memory, not doing math.

Each HBM stack is several DRAM dies connected vertically by through-silicon vias, sitting next to the GPU on a silicon interposer: a thin slab of silicon that carries thousands of short wires between the stacks and the GPU die. In the HBM2 to HBM3e generations each stack presents a 1,024-bit interface, far wider than a single GDDR chip, and the wires are short enough to run at moderate clock speeds with low energy per bit. Bandwidth is interface width times transfer rate, so many wide, short connections add up to terabytes per second. The cost is manufacturing: stacking and interposer packaging are harder and more expensive than soldering GDDR chips around a board, which is why HBM stays on data center parts. It is one way to build GPU RAM.

HBM generations on real GPUs

The names run HBM2, HBM2e, HBM3 and HBM3e (the "e" marks an enhanced, faster refresh), each raising capacity and speed per stack. Datasheet figures for specific GPUs:

GPUMemoryBandwidth
V10016GB or 32GB HBM2900GB/s
A100 (80GB)80GB HBM2e2,039GB/s
H100 SXM80GB HBM33.35TB/s
H200141GB HBM3e4.8TB/s
B200180GB HBM3e8TB/s
MI300X192GB HBM35.3TB/s

Form factor matters: the PCIe H100 uses HBM2e at about 2TB/s, well below the SXM card. Bandwidth rose about 9x from V100 to B200 (8,000 / 900, computed). The H200 keeps the H100's compute but has 76% more memory and 43% more bandwidth. NVIDIA's announced Vera Rubin platform moves to HBM4, with 288GB and 22TB/s per GPU in the figures NVIDIA has disclosed.

Why bandwidth matters for LLM inference

Generating each token means reading essentially every weight of the model from memory once, while doing only about two arithmetic operations per weight. That is very little math per byte moved (low arithmetic intensity), so the compute units mostly wait on memory. For a single request, the fastest possible decode speed is roughly memory bandwidth divided by the bytes of weights.

Take Llama 3.1 8B in FP16: 8 billion parameters x 2 bytes = 16GB of weights.

GPUMemory typeBandwidthCeiling (bandwidth / 16GB)
RTX 4090GDDR6Xabout 1,008GB/sabout 63 tokens/s
H100 SXMHBM33,350GB/sabout 209 tokens/s
H200HBM3e4,800GB/s300 tokens/s
B200HBM3e8,000GB/s500 tokens/s

These are theoretical ceilings for one request, not benchmarks. Real servers also read the KV cache, carry overheads, and batch many requests together. The ordering still follows bandwidth, which is why a memory-bound workload gains on an H200 even though its compute matches the H100.

HBM vs GDDR

GDDR (GDDR6, GDDR6X, GDDR7) uses ordinary memory chips soldered around the GPU, connected over a narrower, faster-clocked bus. It is simpler and less expensive to build, and consumer cards use it: the RTX 4090 has 24GB of GDDR6X, the RTX 5090 has 32GB of GDDR7 at 1,792GB/s, and the L40S has 48GB of GDDR6 at 864GB/s. Next to the H100's 3.35TB/s, that L40S figure is about 3.9x lower (3,350 / 864, computed). GDDR cards are strong for models that fit and for light traffic, and they fall behind on large-batch or long-context serving.

What it means when you pick a GPU

Match the memory type to the job. If the model fits in a GDDR card's VRAM and you serve a few requests at a time, a lower-bandwidth card can be enough, and consumer cards usually rent for less per hour than data center ones.

If you serve many users at once, use long contexts, or run large models, bandwidth and capacity decide it. The H100 (80GB, 3.35TB/s) is the baseline. The H200 adds 141GB and 4.8TB/s on the same compute, which pays off when the KV cache or memory-bound decoding is the bottleneck and buys nothing if the job is limited by arithmetic. The MI300X offers 192GB at 5.3TB/s, the most memory of the cards above, but runs on AMD's ROCm software stack, so confirm your framework supports it. Aquanode rents GPUs by the hour, with current rates on the pricing page.

Building on GPUs? Aquanode runs the workload.

Deploy on H100, H200, B200, A100 and MI300X across a multi-provider marketplace, without racking your own hardware or committing to one cloud's spec sheet.

See also

Submit the job. Everything after that is ours.

Sign up in 60 seconds. Pay for the GPU minutes you actually use.

© 2026 Aquanode. All rights reserved.

All trademarks, logos and brand names are the property of their respective owners.