FlashInfer: GPU Kernels for LLM Inference (2026 Guide)

Back
Team Aquanode

Team Aquanode

Sarthak Vaish

Updated OCTOBER 8, 2026Published OCTOBER 8, 2026

FlashInfer is an open-source library of GPU kernels and a kernel generator for LLM inference. It provides attention (including paged and ragged KV caches), plus GEMM, mixture-of-experts, sampling and communication kernels, and it sits underneath engines such as vLLM and SGLang. You rarely call it yourself: you pick it, or the engine picks it for you, as the attention backend.

This post covers what is in it, how it differs from FlashAttention, how to install it and run the quickstart, which GPUs it supports, and how to select it in vLLM and SGLang. It is part of our guide to LLM inference engines.

TL;DR

  • FlashInfer is a serving-focused kernel library. Its attention kernels are built for paged KV caches, batched decode and shared prefixes, which are the shapes inference engines actually produce.
  • It is used inside the major engines. The project's README lists SGLang, vLLM, TensorRT-LLM, TGI, MLC-LLM, LightLLM, lorax and ScaleLLM among its users.
  • It covers more than attention. The README lists GEMM (BF16, FP8, FP4), fused MoE, sampling and custom all-reduce.
  • Defaults differ by engine and GPU. vLLM lists FlashInfer first for automatic selection on Blackwell, SGLang defaults to FlashInfer on GPUs such as the A100 and Ada cards, and both let you force it with a flag.
  • Verdict: use whichever backend your engine selects, and try the other when you need to debug or squeeze performance on a specific GPU. Measure on your own model, since published kernel speedups are against specific baselines.

What FlashInfer is

The FlashInfer paper describes it as a customizable and efficient attention engine for LLM serving, integrated into SGLang, vLLM and MLC-Engine. The repository today describes a wider scope: a library and kernel generator for inference with unified APIs for attention, GEMM and MoE. The docs at version 0.7.0.post1 (the latest stable release on the project's GitHub releases page when we checked on October 8, 2026) also list communication, sampling, normalization and RoPE kernels.

Three ideas from the paper explain the design:

  1. A block-sparse KV cache format. One format can describe a padded tensor, a ragged tensor or a page table, and composable formats aim to improve memory access and reduce redundancy. This is how it reads a paged KV cache efficiently.
  2. A JIT-customizable attention template. Instead of shipping one fixed kernel, it compiles attention variants on demand, so users can adapt attention behavior to different settings.
  3. Load-balanced scheduling. Requests in a batch have wildly different lengths. FlashInfer plans the work split to even out the load while staying compatible with CUDA graphs, which need static launch configurations.

FlashInfer vs FlashAttention

They are neighbours, not rivals. FlashAttention is the reference fast-attention kernel family and the one most people meet first. FlashInfer targets the serving case: many requests at once, each with a different context length, KV stored in pages, often sharing prefixes.

FlashAttentionFlashInfer
Main useFast exact attention, widely used in training and inferenceLLM serving
KV cache layoutsSee its READMEPadded, ragged and paged, with a block-sparse format
Beyond attentionNoGEMM, MoE, sampling, communication
CustomizationFixed kernels per versionJIT-compiled variants

Engines often use both: a backend per phase, or per GPU. The vLLM docs list FLASH_ATTN and FLASHINFER as separate backends, and SGLang lets you set prefill and decode backends separately.

Features (from the repository README)

  • Attention: paged and ragged KV cache, decode, prefill and append kernels, DeepSeek MLA, cascade attention for shared prefixes, sparse attention, and POD-Attention for mixed batching.
  • GEMM: BF16, FP8 with per-tensor and groupwise scaling, FP4 (NVFP4 and MXFP4), grouped GEMM. See NVFP4 vs MXFP4 and FP8.
  • MoE: fused kernels for models such as DeepSeek-V3 and Llama 4, top-k routing, and FP8 and FP4 experts. See mixture of experts.
  • Sampling: sorting-free top-k, top-p and min-p, plus chain speculative sampling. See speculative decoding.
  • Communication: custom all-reduce, multi-node NVLink, NVSHMEM integration.
  • Other: RoPE, RMSNorm and LayerNorm, fused SiLU and GELU.

Cascade attention for shared prefixes

When many requests share a long prefix, decoding re-reads that shared KV for every request. Cascade attention computes attention over the shared part once and merges it with each request's private part, using a merge-state operation (the API docs list merge_state and merge_states in flashinfer.cascade). FlashInfer's launch post reports up to 31 times speedup over the baseline vLLM PagedAttention implementation for shared-prefix batch decoding, in a test with a 32,768-token prompt and a batch of 256 (their benchmark, against vLLM v0.2.6).

What the project publishes about performance

All of these are the authors' figures, on their setups, and the baselines are named versions that have since changed.

  • Paper: inter-token latency down 29% to 69% against compiler backends on a serving benchmark, 28% to 30% lower latency for long-context inference, and 13% to 17% faster parallel generation.
  • Launch post (February 2024): up to 2 to 3 times speedup for grouped-query attention on A100 and H100 versus vLLM, by running GQA decode on tensor cores, and 3 times faster than vLLM PageAttention for batched GQA decode at batch size 64. The comparisons use FlashAttention 2.4.2 and vLLM v0.2.6.

Treat these as direction, not a promise: both engines have moved on since. For your own numbers, use the method in TTFT and tokens per second.

Install

The README lists three packages: flashinfer-python (the core, which compiles or downloads kernels on first use), flashinfer-cubin (pre-compiled kernel binaries) and flashinfer-jit-cache (architecture-specific pre-built kernels).

pip install flashinfer-python
flashinfer install-cubin-wheel
flashinfer install-jit-cache-wheel
flashinfer show-config

For Blackwell CuTe DSL kernels the README adds an extra:

pip install flashinfer-python[cu13]

If you serve through vLLM or SGLang, check that engine's own install notes first: it may already bundle a compatible FlashInfer build, and mixing versions can break imports.

Quickstart

This single-request decode example is from the README:

import torch
import flashinfer

q = torch.randn(32, 128, device="cuda", dtype=torch.float16)       # [num_qo_heads, head_dim]
k = torch.randn(2048, 32, 128, device="cuda", dtype=torch.float16) # [kv_len, num_kv_heads, head_dim]
v = torch.randn(2048, 32, 128, device="cuda", dtype=torch.float16)

output = flashinfer.single_decode_with_kv_cache(q, k, v)

For batched serving, the API docs describe wrapper classes such as BatchDecodeWithPagedKVCacheWrapper (default KV layout NHD). The pattern in the docs' newer wrappers splits work into a plan step, which takes static geometry like batch size, head counts and dtypes and should run before CUDA graph capture, and a run step, which takes the per-launch tensors such as Q, the KV cache and page tables. Replanning invalidates graphs captured earlier. The docs say the exact signature differs by wrapper, so read the page for the one you use.

Supported GPUs

Per the README (by compute capability): Turing 7.5 (T4, RTX 20 series), Ampere 8.0 and 8.6 (A100, A10, RTX 30 series), Ada 8.9 (L4, L40, RTX 40 series), Hopper 9.0 (H100, H200), Blackwell 10.0 and 10.3 (B200, B300), Blackwell 11.0 (Jetson Thor) and Blackwell 12.0 and 12.1 (RTX 50 series, DGX Spark). The README notes that not every feature is available on every architecture, so check the feature you need against your card.

That range is wider than FlashAttention 3, which is Hopper only. See the H100 and B200 pages for the cards most people run it on.

Using it in vLLM and SGLang

vLLM. Select the backend with --attention-backend. The docs list FLASHINFER with native support on compute capability 8.x to 9.x, XQA on 9.0, and trtllm-gen kernels on 10.x. In automatic selection on Blackwell the order is FLASHINFER first, then FLASH_ATTN; on Ampere and Hopper it is FLASH_ATTN first, then FLASHINFER.

vllm serve <model> --attention-backend FLASHINFER

SGLang. The flag takes lowercase names. Per the docs, Hopper defaults to fa3, Blackwell to trtllm_mha for standard attention (with flashinfer for MLA), and other GPUs such as the A100 or Ada cards to flashinfer if it is installed, otherwise triton. SGLang also lets you set prefill and decode backends separately.

python3 -m sglang.launch_server --model-path <model> --attention-backend flashinfer

SGLang's docs call auto-selection best-effort, and defaults in both engines change between versions, so pin the backend explicitly when you benchmark. For how these engines compare, read vLLM vs TensorRT-LLM vs SGLang and serving LLMs with vLLM.

When FlashInfer is worth trying

  • You serve many concurrent requests with very different context lengths.
  • Many requests share a long system prompt or document, so cascade attention applies.
  • You are on a GPU generation your engine's default does not cover well, such as new Blackwell parts.
  • You use quantized GEMM or MoE on Blackwell and want the FP4 kernels the library ships.

If your server already runs well on the engine's default, there is no reason to change it for its own sake.

Run it on a cloud GPU

To compare backends, run the same model on a Hopper card and a Blackwell card and pin each backend in turn. The boxes below show what is available to rent right now.

FAQ

Is FlashInfer faster than FlashAttention?

It depends on the shape of the work. The project reports large gains for shared-prefix and grouped-query decode against the vLLM baseline of early 2024. On large dense batches the gap narrows, and both are updated often. Benchmark your model.

Do I need to install FlashInfer to use vLLM or SGLang?

Usually not for the default path. Engines choose a backend automatically, and FlashInfer is used when it is the selected backend. Install it yourself for direct use in your own code.

Does FlashInfer work on consumer GPUs?

The README lists Ada (RTX 40 series) and Blackwell 12.x (RTX 50 series, DGX Spark) among supported architectures, as well as RTX 30 and 20 series cards. Feature support varies by architecture.

Is FlashInfer only for attention?

No. The repository also provides GEMM, MoE, sampling, communication and normalization kernels.

What is the difference between flashinfer-python, flashinfer-cubin and flashinfer-jit-cache?

The core package compiles or downloads kernels on first use. The other two ship pre-built kernels so the first call does not pay a compile cost, with one specific to your GPU architecture.

Which version is current?

When we checked the GitHub releases page on October 8, 2026, the latest stable tag was v0.7.0.post1, with a v0.7.1 release candidate listed. Check the releases page before pinning.

Sources

#llm inference#inference engines#flashinfer#attention kernels#paged kv cache#blackwell

Submit the job. Everything after that is ours.

Sign up in 60 seconds. Pay for the GPU minutes you actually use.

© 2026 Aquanode. All rights reserved.

All trademarks, logos and brand names are the property of their respective owners.