FlashAttention 3 is a rewrite of the attention kernel for NVIDIA Hopper GPUs (H100, H200) that its authors report is 1.5 to 2.0 times faster than FlashAttention 2 on the H100, reaching up to 740 TFLOPS in FP16, about 75% of the chip's peak. FlashAttention 2 is the version that runs on a wider range of GPUs (Ampere, Ada and Hopper), so on an A100 or an RTX 4090 it is the one you get.
For most people the choice is made by the inference engine, not by you. This post explains what each version changed, what the papers measured, the GPU support limits, and how to check or force which one your server is using. It is part of our guide to LLM inference engines.
TL;DR
- FlashAttention 2 (Tri Dao, 2023) reaches 50% to 73% of theoretical peak on an A100, about twice the original FlashAttention. It supports Ampere, Ada and Hopper.
- FlashAttention 3 (Tri Dao and co-authors, 2024) uses Hopper-only hardware features to lift H100 utilization from about 35% to about 75% in FP16 and adds an FP8 path near 1.2 PFLOPS. It needs an H100 or H800 class GPU.
- FlashAttention 4 targets Blackwell (and Hopper) and is now in the same repository. The paper reports up to 1,613 TFLOPS on a B200.
- Verdict: on Hopper use FA3 (vLLM and SGLang already default to it). On Ampere and Ada use FA2. On Blackwell, use whatever your engine selects, which is FA4 or FlashInfer depending on the engine. You only need to choose by hand to debug.
Why attention kernels matter
Attention compares every token's query with the keys of the tokens it can see. Written naively, the score matrix is as large as sequence length squared, and writing it to GPU memory and reading it back is slow. The original FlashAttention (Dao et al., 2022) made the algorithm IO-aware: it tiles the computation so intermediate results stay in fast on-chip memory, still computing exact attention. The authors reported about a 15% end-to-end speedup on BERT-large, roughly 3 times on GPT-2 and 2.4 times on the Long-Range Arena benchmark.
Because the kernel runs in every layer on every token, small gains compound. It is also why the choice of attention backend appears in every inference engine, as covered in serving LLMs with vLLM. The glossary entry on flash attention has the short version.
FlashAttention 2: better use of the GPU you already have
The FlashAttention 2 paper observed that the first version reached only 25% to 40% of an A100's theoretical peak, far below what a well-tuned matrix multiply reaches (the paper cites 80% to 90%). Its three changes, per the abstract:
- Fewer non-matmul operations. Tensor cores are much faster than the units doing softmax bookkeeping, so the algorithm was tweaked to spend less time off the tensor cores.
- More parallelism. Work is split across thread blocks along the sequence length, even for a single head, which raises occupancy when batch size or head count is small (the long-context case).
- Better work partitioning inside a block. Work is divided among warps to cut shared-memory communication.
Published result: about 2 times faster than FlashAttention, 50% to 73% of A100 theoretical peak FLOPs, and up to 225 TFLOPS per A100 in end-to-end GPT-style training, which the paper states as 72% model FLOPs utilization. (Computed from those two figures: that implies an A100 peak of about 312 TFLOPS.)
Support (from the repository README): CUDA 12.0 or newer and PyTorch 2.2 or newer on Linux; Ampere, Ada or Hopper GPUs (examples named: A100, RTX 3090, RTX 4090, H100); fp16 and bf16, with bf16 needing Ampere or newer; head dimensions up to 256. Turing is in a separate repository.
FlashAttention 3: built for Hopper
By 2024 the H100 had hardware features that FA2 did not use, and the FA3 paper reports FA2 reaching only 35% utilization on it. FA3 uses three techniques:
- Overlap data movement and compute. Warp specialization lets some warps issue asynchronous loads (through the Tensor Memory Accelerator, TMA) while others feed the asynchronous tensor cores.
- Interleave matmul and softmax. Softmax's exponential runs on units far slower than the tensor cores: the write-up gives 3.9 TFLOPS for the exponential versus 989 TFLOPS for FP16 matmul on an H100 SXM5. Overlapping the two hides it. In the authors' breakdown (FP16 forward, head dimension 128, 8K sequence), the overlap tricks lift throughput from about 570 to 620 TFLOPS with pingpong scheduling, and further to about 640 to 660 TFLOPS with intra-warpgroup pipelining.
- FP8 with lower error. Block quantization and incoherent processing (a random Hadamard transform on queries and keys) cut FP8 error by 2.6 times versus a baseline FP8 attention in the authors' test.
Published results, all from the authors on H100:
| Measure | FlashAttention 2 | FlashAttention 3 |
|---|---|---|
| H100 utilization (FP16) | about 35% | up to 75% |
| FP16 forward throughput | not stated in the paper abstract | up to 740 TFLOPS |
| Speedup | baseline | 1.5 to 2.0 times over FA2 (the write-up's benchmark section gives 1.6 to 2.0) |
| FP8 | not available | close to 1.2 PFLOPS |
Computed sanity check: 740 divided by 989 is about 75%, matching the stated utilization, and 35% of 989 is about 346 TFLOPS for FA2 at that utilization. Your own speedup depends on sequence length, head dimension and batch shape, so expect the low end of the range at short sequences.
Support (from the repository README): FA3 is labelled beta. It needs an H100 or H800, CUDA 12.3 or newer (12.8 recommended), and ships FP16 and BF16 forward and backward plus FP8 forward. Install from the repository's hopper directory:
cd hopper
python setup.py install
export PYTHONPATH=$PWD
pytest -q -s test_flash_attn.py
from flash_attn_3 import flash_attn_interface
flash_attn_interface.flash_attn_func()
FA2 installs with pip, per the same README:
pip install flash-attn --no-build-isolation
On a low-RAM machine the README suggests limiting parallel compile jobs with MAX_JOBS=4.
And FlashAttention 4
The same repository now has FlashAttention 4, written in CuTe-DSL and aimed at Hopper and Blackwell (H100, B200). The paper (Zadouri, Hoehnerbach, Shah, Liu, Thakkar and Dao, March 2026) reports, on B200 with BF16, up to 1.3 times the speed of cuDNN 9.13 and 2.7 times Triton, reaching up to 1,613 TFLOPS at 71% utilization. The reason a new version was needed, per the abstract, is that Blackwell scales tensor core throughput faster than shared memory bandwidth and exponential units, so the bottlenecks moved again. Install per the README: pip install flash-attn-4.
Side by side
| FA2 | FA3 | FA4 | |
|---|---|---|---|
| Target GPUs | Ampere, Ada, Hopper | Hopper (H100, H800) | Hopper and Blackwell |
| Precision | fp16, bf16 | fp16, bf16 forward and backward, fp8 forward | See the repository |
| Reported peak (authors) | 225 TFLOPS in training on A100 | 740 TFLOPS FP16, about 1.2 PFLOPS FP8 on H100 | 1,613 TFLOPS BF16 on B200 |
| Status in the README | stable | beta | newest |
Which version your engine uses
vLLM. The FLASH_ATTN backend picks the version by GPU, per the vLLM attention backend docs: FA4 on Blackwell (compute capability 10 and up), FA3 on Hopper (9.x), FA2 otherwise (8.0 and up). You can set the version with --attention-config.flash_attn_version=2, 3 or 4, and the backend with --attention-backend. For automatic selection on Blackwell, the same page lists FlashInfer first, then FlashAttention, so your Blackwell server may not run FA4 unless you ask for it. See FlashInfer for that alternative.
vllm serve <model> --attention-backend FLASH_ATTN
SGLang. Per its attention backend docs, Hopper defaults to fa3 for standard attention (with CUDA 12.3 or newer), Blackwell defaults to trtllm_mha unless speculative decoding uses top-k above 1, and other GPUs default to flashinfer when available, otherwise triton. Force a backend with --attention-backend. The docs also list fa4. The docs note that auto-selection is best-effort.
python3 -m sglang.launch_server --model-path <model> --attention-backend fa3
Engine comparisons that depend on these kernels are in vLLM vs TensorRT-LLM vs SGLang.
Do FA2 and FA3 give the same answer?
Both compute exact attention, so for the same inputs the outputs match up to normal floating-point differences. The exception is FA3's FP8 path, which trades precision for speed and is the part to validate on your own evaluations before you rely on it.
When to choose by hand
- Debugging. If a model crashes or gives bad output on one backend, switch to another (including the Triton or PyTorch fallbacks) to see if the kernel is the cause.
- Unusual head dimensions or positional schemes. Not every kernel supports every head size or attention variant. ALiBi models, for example, cannot use a kernel built for rotary positions.
- Benchmarking. Pin the version so a library upgrade does not change your baseline.
- Building from source. FA3 and FA4 install separately from
flash-attn, so a barepip install flash-attngives you FA2.
Run it on a cloud GPU
The FA2 versus FA3 difference only shows up on Hopper, so that is where to test it. Rent an H100 or H200, pin each backend in turn, and compare time to first token and tokens per second with the method in TTFT and tokens per second.
FAQ
Is FlashAttention 3 faster than FlashAttention 2 on an A100?
FA3 does not run on an A100. It needs Hopper. On Ampere and Ada the choice is FA2.
Do I have to install FlashAttention separately to use vLLM?
Not normally. vLLM ships its own build of the FlashAttention kernels and selects the version by GPU. Installing the standalone package is for your own code, such as a training loop or a custom model.
Does FlashAttention 3 change the model's output?
In FP16 and BF16 it computes the same exact attention as FA2. The FP8 path introduces quantization error, which the authors bound at 2.6 times lower than a baseline FP8 attention, but you should still evaluate it.
Is FlashAttention only for long context?
It helps at any length, and the gap between FA2 and FA3 is larger at longer sequences. The memory savings matter most when sequences are long, because the score matrix is never written out in full.
Should I use FlashAttention or FlashInfer?
They overlap. FlashInfer is built around serving: paged KV caches, batched decode and cascade attention. See the FlashInfer guide. On Hopper, vLLM's automatic choice is FlashAttention; on Blackwell, its automatic order lists FlashInfer first.
Sources
- Dao et al., FlashAttention: https://arxiv.org/abs/2205.14135
- Dao, FlashAttention-2: https://arxiv.org/abs/2307.08691 (details confirmed at https://ar5iv.labs.arxiv.org/html/2307.08691)
- FlashAttention-3 paper: https://arxiv.org/abs/2407.08608
- Tri Dao, FlashAttention-3 write-up: https://tridao.me/blog/2024/flash3/
- Zadouri et al., FlashAttention-4: https://arxiv.org/abs/2603.05451
- FlashAttention repository and README (support, install, FA2 FA3 FA4): https://github.com/Dao-AILab/flash-attention
- vLLM attention backends: https://docs.vllm.ai/en/latest/design/attention_backends.html
- SGLang attention backends: https://docs.sglang.io/advanced_features/attention_backend.html