What is HBM (High Bandwidth Memory)?
Abbreviated HBM
HBM (High Bandwidth Memory) is DRAM built as vertical stacks of memory chips and mounted in the same package as the GPU die, which gives it far more bandwidth than the GDDR memory on consumer cards. Data center GPUs such as the H100, H200 and MI300X use it because LLM inference spends most of its time waiting on memory, not doing math.
Each HBM stack is several DRAM dies connected vertically by through-silicon vias, sitting next to the GPU on a silicon interposer: a thin slab of silicon that carries thousands of short wires between the stacks and the GPU die. In the HBM2 to HBM3e generations each stack presents a 1,024-bit interface, far wider than a single GDDR chip, and the wires are short enough to run at moderate clock speeds with low energy per bit. Bandwidth is interface width times transfer rate, so many wide, short connections add up to terabytes per second. The cost is manufacturing: stacking and interposer packaging are harder and more expensive than soldering GDDR chips around a board, which is why HBM stays on data center parts. It is one way to build GPU RAM.
HBM generations on real GPUs
The names run HBM2, HBM2e, HBM3 and HBM3e (the "e" marks an enhanced, faster refresh), each raising capacity and speed per stack. Datasheet figures for specific GPUs:
| GPU | Memory | Bandwidth |
|---|---|---|
| V100 | 16GB or 32GB HBM2 | 900GB/s |
| A100 (80GB) | 80GB HBM2e | 2,039GB/s |
| H100 SXM | 80GB HBM3 | 3.35TB/s |
| H200 | 141GB HBM3e | 4.8TB/s |
| B200 | 180GB HBM3e | 8TB/s |
| MI300X | 192GB HBM3 | 5.3TB/s |
Form factor matters: the PCIe H100 uses HBM2e at about 2TB/s, well below the SXM card. Bandwidth rose about 9x from V100 to B200 (8,000 / 900, computed). The H200 keeps the H100's compute but has 76% more memory and 43% more bandwidth. NVIDIA's announced Vera Rubin platform moves to HBM4, with 288GB and 22TB/s per GPU in the figures NVIDIA has disclosed.
Why bandwidth matters for LLM inference
Generating each token means reading essentially every weight of the model from memory once, while doing only about two arithmetic operations per weight. That is very little math per byte moved (low arithmetic intensity), so the compute units mostly wait on memory. For a single request, the fastest possible decode speed is roughly memory bandwidth divided by the bytes of weights.
Take Llama 3.1 8B in FP16: 8 billion parameters x 2 bytes = 16GB of weights.
| GPU | Memory type | Bandwidth | Ceiling (bandwidth / 16GB) |
|---|---|---|---|
| RTX 4090 | GDDR6X | about 1,008GB/s | about 63 tokens/s |
| H100 SXM | HBM3 | 3,350GB/s | about 209 tokens/s |
| H200 | HBM3e | 4,800GB/s | 300 tokens/s |
| B200 | HBM3e | 8,000GB/s | 500 tokens/s |
These are theoretical ceilings for one request, not benchmarks. Real servers also read the KV cache, carry overheads, and batch many requests together. The ordering still follows bandwidth, which is why a memory-bound workload gains on an H200 even though its compute matches the H100.
HBM vs GDDR
GDDR (GDDR6, GDDR6X, GDDR7) uses ordinary memory chips soldered around the GPU, connected over a narrower, faster-clocked bus. It is simpler and less expensive to build, and consumer cards use it: the RTX 4090 has 24GB of GDDR6X, the RTX 5090 has 32GB of GDDR7 at 1,792GB/s, and the L40S has 48GB of GDDR6 at 864GB/s. Next to the H100's 3.35TB/s, that L40S figure is about 3.9x lower (3,350 / 864, computed). GDDR cards are strong for models that fit and for light traffic, and they fall behind on large-batch or long-context serving.
What it means when you pick a GPU
Match the memory type to the job. If the model fits in a GDDR card's VRAM and you serve a few requests at a time, a lower-bandwidth card can be enough, and consumer cards usually rent for less per hour than data center ones.
If you serve many users at once, use long contexts, or run large models, bandwidth and capacity decide it. The H100 (80GB, 3.35TB/s) is the baseline. The H200 adds 141GB and 4.8TB/s on the same compute, which pays off when the KV cache or memory-bound decoding is the bottleneck and buys nothing if the job is limited by arithmetic. The MI300X offers 192GB at 5.3TB/s, the most memory of the cards above, but runs on AMD's ROCm software stack, so confirm your framework supports it. Aquanode rents GPUs by the hour, with current rates on the pricing page.
Building on GPUs? Aquanode runs the workload.
Deploy on H100, H200, B200, A100 and MI300X across a multi-provider marketplace, without racking your own hardware or committing to one cloud's spec sheet.
See also
VRAM
VRAM is the memory attached to a GPU that holds the data it works on, and it caps which AI models fit. VRAM vs RAM, how to check yours, and how much AI needs.
GPU RAM
GPU RAM is the large off-die memory pool every Streaming Multiprocessor shares, built from slower, denser DRAM cells rather than the SRAM used in registers and cache.
Arithmetic Intensity
Arithmetic intensity is the ratio of compute operations to bytes moved in a kernel. Why it decides whether a workload is compute-bound or memory-bound, and how tricks like recomputation trade memory traffic for extra FLOPs.
KV Cache
A KV cache stores the key and value tensors of past tokens so an LLM never recomputes them. It grows with context length and batch size, and it eats VRAM.
FlashAttention
FlashAttention is an exact attention algorithm that tiles the computation so the full attention matrix is never written to GPU memory, cutting HBM traffic.
TFLOPS
TFLOPS means trillions of floating-point operations per second, a GPU's peak math speed. Why precision changes the number and why it won't predict LLM speed.