What is FlashAttention?
FlashAttention is an exact attention algorithm that computes transformer attention in small tiles inside fast on-chip memory, so the full attention matrix is never written out to the GPU's main memory. Because it returns the same result as standard attention (it is not an approximation), it works as a drop-in replacement that cuts reads and writes to HBM, runs faster, and lets models handle much longer sequences.
It was introduced in the paper "FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness" (Dao et al., 2022).
The problem it fixes
Attention multiplies queries by keys to get a score for every pair of tokens, applies a softmax, then multiplies by the values. A straightforward implementation runs these as separate GPU kernels, and each kernel writes its result to HBM for the next one to read back. The score matrix has one entry per pair of tokens, so it grows with the square of sequence length.
For one sequence of 32,768 tokens in 16-bit, one head in one layer produces 32,768 x 32,768 x 2 bytes = 2,147,483,648 bytes, which is 2 GiB of scores. The arithmetic on those scores is light, so the GPU spends its time moving bytes rather than computing: a textbook case of being memory-bound on the roofline model.
How FlashAttention works
It splits the queries, keys and values into blocks sized to fit in on-chip SRAM, which on a GPU means shared memory and registers inside each streaming multiprocessor. For each block it computes the scores, folds them into a running softmax (using an online softmax technique that tracks a running maximum and sum, so the exact answer can be assembled tile by tile), and accumulates the output. All of this is fused into one kernel, so the intermediate scores never leave the chip. In the backward pass during training, it recomputes scores instead of storing them, trading a little extra arithmetic for far less memory traffic.
The result is that the extra memory attention needs grows linearly with sequence length instead of quadratically. The total arithmetic is not reduced; the win comes from moving fewer bytes. Later versions of FlashAttention from the same research line improve how work is split across the GPU and take advantage of newer hardware features.
One thing it does not do is shrink the KV cache. During token-by-token generation each new token still reads the whole cache from HBM, so decoding stays limited by memory bandwidth.
What it means when you pick a GPU
Memory bandwidth still matters. FlashAttention removes the quadratic score traffic, but weights and the KV cache still stream from HBM on every generated token. These are the bandwidth figures for the cards we list:
| GPU | Memory | Bandwidth |
|---|---|---|
| H100 | 80GB HBM3 | 3.35 TB/s |
| H200 | 141GB HBM3e | 4.8 TB/s |
| B200 | 180GB HBM3e | 8 TB/s |
| AMD MI300X | 192GB HBM3 | 5.3 TB/s |
| RTX 4090 | 24GB GDDR6X | 1,008 GB/s (derived from bus width and transfer rate) |
On-chip memory sets tile size. Each streaming multiprocessor has a limited amount of shared memory and registers, and that limit decides how big a tile can be and how often the kernel has to go back to HBM. Check this on a card's spec sheet rather than assuming two GPUs behave alike.
Support depends on GPU generation. The fastest FlashAttention kernels are tuned for particular GPU generations. An older card may be limited to an earlier version or fall back to a slower attention implementation, and AMD cards go through a separate software stack (see ROCm vs CUDA). Before you rent, confirm that your framework's attention backend actually supports the card.
Context length is then limited by the KV cache. With FlashAttention in place, the thing that usually runs out first on a long prompt is VRAM for the cache, not the attention matrix. Estimate how much you need with the VRAM calculator before you pick a card. Aquanode rents GPUs by the hour.
Building on GPUs? Aquanode runs the workload.
Deploy on H100, H200, B200, A100 and MI300X across a multi-provider marketplace, without racking your own hardware or committing to one cloud's spec sheet.
See also
KV Cache
A KV cache stores the key and value tensors of past tokens so an LLM never recomputes them. It grows with context length and batch size, and it eats VRAM.
HBM (High Bandwidth Memory)
HBM is stacked DRAM packaged beside a GPU die, giving data center cards several TB/s of memory bandwidth. Why LLM inference depends on it, and HBM vs GDDR.
Shared Memory
Shared memory is the fast, on-chip pool of memory a CUDA thread block uses to avoid repeatedly hitting slower global memory. The standard load-compute-store pattern it enables, and where bank conflicts come from.
Roofline Model
The roofline model plots a kernel's arithmetic intensity against two hardware ceilings, memory bandwidth and arithmetic bandwidth, to show at a glance whether it's compute-bound or memory-bound. Where it came from and why GPUs need it.
Tensor Core
A Tensor Core is the GPU hardware unit that executes an entire matrix multiply-accumulate as one instruction instead of one scalar multiply at a time. How that trade unlocks NVIDIA's highest FLOP counts, and why an H100 has only four of them per SM.
BF16 (bfloat16)
BF16 is a 16-bit float with FP32's 8 exponent bits but only 7 mantissa bits. It is the default for training and costs 2 bytes per model parameter.