NVIDIA Transformer Engine: FP8 Training Guide (2026)

Back
Team Aquanode

Team Aquanode

Sarthak Vaish

Updated OCTOBER 8, 2026Published OCTOBER 8, 2026

NVIDIA Transformer Engine is an open-source library that makes Transformer models train and run in FP8 on Hopper and newer GPUs by handling the hard parts for you: choosing FP8 formats, tracking the scaling factors that keep values inside FP8's tiny range, and fusing the kernels. You wrap your layers in its modules, turn on an autocast context with a recipe, and the library manages the low-precision details while master weights and sensitive operations stay in higher precision.

This guide covers:

  • What Transformer Engine is and what it replaces in your training code
  • How the two FP8 formats work and why scaling is needed
  • The scaling recipes (delayed, current, block, MXFP8, NVFP4) and which hardware supports each
  • A minimal PyTorch example
  • NVIDIA's published claims, labelled as the vendor's, and how to think about cost

TL;DR

  • Transformer Engine (TE) is NVIDIA's library for FP8 Transformer training and inference on PyTorch and JAX. FP8 is supported on Hopper, Ada, Blackwell and Rubin GPUs; MXFP8 and NVFP4 training are Blackwell features, per the project README.
  • FP8 comes in two formats: E4M3 (more precision, max value 448) and E5M2 (more range, max value 57,344). The default HYBRID format uses E4M3 for the forward pass and E5M2 for gradients.
  • Because FP8 has so little range, TE keeps a scaling factor per tensor (or per block) and updates it from the observed maximum value, called the amax.
  • Verdict: if you train or fine-tune Transformers on H100 or H200 and you are still in BF16, FP8 through TE is the first optimization to try. Validate loss against a BF16 run before trusting it.

For the full list of datacenter accelerators and how they fit together, see our datacenter GPU overview.

What is Transformer Engine?

The project README describes it as an NVIDIA library for accelerating Transformer model training on NVIDIA GPUs. It provides:

  • Drop-in modules for PyTorch and JAX, such as linear layers, layer-norm-plus-linear layers, MLP blocks and full Transformer layers, written to run their matrix multiplies on tensor cores in low precision.
  • Scaling management. TE tracks the scaling factors, amax histories and quantization metadata that FP8 needs, so you do not write that bookkeeping yourself.
  • Fused kernels that combine operations (casting, normalization, attention pieces) to avoid extra trips through memory.

NVIDIA introduced the idea with the Hopper architecture. The Hopper in-depth blog describes the H100's transformer engine as software plus custom Hopper tensor-core technology that "intelligently manages and dynamically chooses between FP8 and 16-bit calculations", analyzing the statistics of each layer's output and scaling data into the representable range so that every layer "operates with exactly the range it requires."

The key point for practitioners: FP8 is not a flag you flip on a PyTorch model. Done naively, FP8 overflows or underflows and training diverges. TE exists to make it safe and fast.

Which GPUs are supported?

From the README:

PrecisionSupported on
FP8 trainingHopper, Ada, Blackwell and Rubin GPUs
MXFP8 and NVFP4 trainingBlackwell GPUs
FP16 and BF16 optimizationsAmpere and later GPUs

The same page lists release news with dates: version 2.17 in July 2026, 2.18 in August 2026 (FP8 block scaling in PyTorch), and 2.19 in September 2026 (Rubin support, hybrid quantization and expanded FP8 attention support). Check the release notes for the version you install, since recipe support has moved quickly.

How FP8 works

Two formats: E4M3 and E5M2

NVIDIA's Transformer Engine documentation describes the two 8-bit formats:

FormatBits (sign / exponent / mantissa)Largest valueTrade-off
E4M31 / 4 / 3448 (plus NaN)More precision, less range
E5M21 / 5 / 257,344 (plus infinity and NaN)More range, less precision

Neither format can represent a typical tensor without help. Activations and gradients span many orders of magnitude, while E4M3 tops out at 448.

TE's Format.HYBRID option uses E4M3 during the forward pass and E5M2 during the backward pass. The documentation's reasoning is that forward activations and weights need precision, while gradients need dynamic range. HYBRID is the format used in NVIDIA's own example.

For a plain-language definition of the type, see FP8 in our glossary, and for how tensor cores execute it, tensor cores.

Why scaling factors are needed

Before a tensor is cast to FP8, it is multiplied by a scaling factor so its largest value lands near the top of the FP8 range. The result is stored with that factor, and the factor is divided back out after the matrix multiply. The scale comes from the tensor's amax. Different recipes differ in how and when they compute it.

The scaling recipes

TE's transformer_engine.common.recipe module offers several. The documentation lists these classes: DelayedScaling, Float8CurrentScaling, Float8BlockScaling, MXFP8BlockScaling, NVFP4BlockScaling and CustomRecipe.

Delayed scaling

Each module keeps a rolling history of recent amax values (the length is configurable) and derives the scale from that history, by default the max over the window. Because the scale comes from earlier iterations, TE does not need to read the tensor twice, which the documentation presents as the efficiency argument: one read per quantization instead of two.

The cost is staleness. NVIDIA's technical blog on scaling strategies notes that the approach assumes tensor value distributions stay fairly stable, and that outliers can dominate the amax history, causing underflow or overflow that may destabilize large runs.

Current scaling

Float8CurrentScaling computes the amax from the live tensor, then casts. That is two passes over the tensor, which the docs note is a significant overhead compared with recipes that need a single read. In exchange it adapts immediately and needs no history buffers. NVIDIA's blog says it is more robust to outliers within the current batch, but it cannot smooth spikes across batches the way a history can. The docs describe it as the simplest recipe and list SM89 (Ada) or newer as the requirement.

Block scaling

Float8BlockScaling divides each tensor into blocks, such as 1 by 128 or 128 by 128, each with its own FP32 scale. Smaller blocks reduce quantization error but increase storage for the scales; larger blocks do the opposite.

MXFP8 and NVFP4 (Blackwell)

MXFP8BlockScaling uses fixed 32-value blocks with E8M0 (power-of-two) scale factors and, per NVIDIA's blog, is a hardware-level solution on Blackwell. NVFP4BlockScaling extends the idea to 4-bit values. These need Blackwell GPUs; the format differences are covered in NVFP4 vs MXFP4, and the hardware in the B200 guide.

Which should you pick?

On Hopper, the practical choices are delayed scaling, current scaling and, in recent versions, FP8 block scaling in PyTorch. A reasonable approach, our suggestion rather than an NVIDIA rule: start with current scaling when you want fewer precision surprises, move to delayed scaling if the extra pass over the tensor shows up in profiling and your loss curve stays matched to BF16, and test block scaling when you see instability from outliers. Always keep a BF16 reference run to compare against.

A minimal PyTorch example

This follows the structure of the FP8 primer in NVIDIA's documentation. The recipe settings (HYBRID format, amax history of 16, max algorithm) are the ones used in that primer.

import torch
import transformer_engine.pytorch as te
from transformer_engine.common.recipe import Format, DelayedScaling

fp8_recipe = DelayedScaling(
    fp8_format=Format.HYBRID,
    amax_history_len=16,
    amax_compute_algo="max",
)

my_linear = te.Linear(768, 768, bias=True)
inp = torch.rand((1024, 768), device="cuda")

with te.autocast(enabled=True, recipe=fp8_recipe):
    out_fp8 = my_linear(inp)

loss = out_fp8.sum()
loss.backward()

Points worth knowing:

  • Backward outside the context. The documentation states that, because of how scaling metadata is aggregated, the backward call needs to happen outside of the autocast context manager. It adds that this has no impact on precision: the backward pass precision follows the forward pass.
  • API naming. Current documentation lists autocast() as the PyTorch context manager and shows fp8_autocast() under deprecated functions. Older tutorials use the old name, so match the name to your installed version.
  • Install. NVIDIA's README gives pip install --no-build-isolation "transformer_engine[pytorch]" (or [jax]), and ready-to-run NGC PyTorch and JAX containers.
  • Availability checks. The API includes functions such as is_fp8_available() and is_mxfp8_available() so a script can test what the GPU supports before choosing a recipe.

For real models you replace te.Linear with TE's larger building blocks or use a framework that already integrates TE, instead of rewriting layers by hand.

Performance: what NVIDIA publishes

All figures are vendor claims with the conditions NVIDIA states. We have not measured them.

  • Hopper launch claims. NVIDIA's Hopper architecture blog claims up to 9x faster AI training and up to 30x faster inference on large language models compared with A100, and says the performance numbers were preliminary and based on current expectations when published.
  • H100 product page. NVIDIA claims up to 4x higher GPT-3 175B training performance, attributing it to a Transformer Engine with FP8, with a footnote reading "Projected performance subject to change". The comparison is an A100 cluster on HDR InfiniBand against an H100 cluster on NDR InfiniBand, so networking differs between the two sides.
  • Peak math. NVIDIA's Hopper write-up lists the H100 SXM5 at 2,000 TFLOPS dense FP8 (4,000 with sparsity). Our arithmetic from NVIDIA's BF16 row on the product page (1,979 TFLOPS with sparsity, so about 990 dense) is that FP8's peak is roughly twice BF16's. Peak is a ceiling; real speedups are lower because of memory-bound operations, casts and communication.
  • Accuracy. NVIDIA's scaling-strategies blog shows validation curves for MXFP8 against BF16 on Nemotron 2B and 8B models tracking closely, and a Nemotron 8B comparison of current-scaling FP8 against BF16 that stays within a narrow margin. The blog shows curves without numeric speedups, so it supports "FP8 can match BF16 quality in those cases", not a speed figure.

Anything beyond that, such as a specific speedup on your model, you should measure. The honest range runs from a modest gain on small models to a larger one on big matmul-dominated Transformers.

Infrastructure and practical needs

  • GPU memory. An 8-bit value takes half the space of a 16-bit one, but a training run also holds other state (optimizer state and higher-precision values for sensitive operations), so do not expect total memory to halve. Measure it on your model. See quantization for the broader picture and BF16 for the format you are likely moving from.
  • Interconnect. FP8 raises compute throughput, which makes communication a bigger share of step time. On multi-GPU jobs NVLink bandwidth matters; the form-factor differences are in SXM vs NVL vs PCIe.
  • Software. Use a recent NGC container or a pinned TE version, and read the release notes: recipe availability changed between 2.17, 2.18 and 2.19 according to the project news.
  • Validation. Run a short BF16 baseline and compare loss curves. A divergence in the first few thousand steps usually points at recipe settings, not hardware.

When to use TE and FP8

Use it when:

  • You train or fine-tune Transformer models on Hopper or Blackwell and your step time is dominated by large matrix multiplies.
  • You serve large models and want FP8 tensor-core throughput. For inference, weight and activation quantization is covered in the quantization glossary.
  • You want the same code path to move from H100 and H200 to Blackwell. TE's README lists the support progression from FP8 on Hopper to MXFP8 and NVFP4 on Blackwell.

Be careful when:

  • Your model has extreme activation outliers. Test current or block scaling instead of delayed scaling.
  • You are on Ampere or older. FP8 tensor cores do not exist there, and TE only offers FP16 and BF16 optimizations.
  • You need bit-exact reproducibility against an old BF16 run. Low precision changes numerics.

Cost: how to compare

We do not type hourly prices here; the live box below shows them. Compare BF16 and FP8 runs with measured throughput: tokens per second times 3,600 gives tokens per GPU-hour on your model. Then divide the live hourly price by that number for cost per token or per training step. If FP8 gives you a measured gain of any size, the cost per token falls by the same ratio at a fixed hourly price, before accounting for any extra steps needed to reach the same loss.

Rent today

Aquanode manages and optimizes GPUs for training and inference workloads, and you can rent the GPUs in the box below on demand. Hopper is the generation where FP8 first became available, so these two are the natural place to try TE.

Specs and availability are on the H100 and H200 pages, with billing at /pricing.

What's next

The TE project news lists Rubin support arriving in version 2.19 in September 2026, alongside MXFP8 and NVFP4 recipes already supported on Blackwell. For what Rubin is and its status, read the Rubin guide, and for the generational picture, Rubin vs Blackwell vs Hopper.

FAQ

What is NVIDIA Transformer Engine?

An open-source NVIDIA library for accelerating Transformer models on NVIDIA GPUs. It provides PyTorch and JAX modules, manages FP8 scaling factors and amax histories, and supplies fused kernels.

Which GPUs support FP8 in Transformer Engine?

Hopper, Ada, Blackwell and Rubin, per the project README. MXFP8 and NVFP4 training are supported on Blackwell. FP16 and BF16 optimizations apply from Ampere onward.

What is the difference between E4M3 and E5M2?

E4M3 has 4 exponent and 3 mantissa bits and a maximum of 448, so it is more precise. E5M2 has 5 exponent and 2 mantissa bits and a maximum of 57,344, so it has wider range. HYBRID uses E4M3 forward and E5M2 backward.

Does FP8 training hurt accuracy?

NVIDIA's published validation curves show MXFP8 and current-scaling FP8 tracking BF16 closely on the Nemotron models it tested. That does not guarantee your model will; compare against a BF16 baseline.

Is delayed scaling or current scaling better?

Delayed scaling reads each tensor once but depends on a history that outliers can distort. Current scaling reads twice but adapts to the live tensor. Start with current scaling if stability matters, and profile before switching.

Do I need Transformer Engine to use FP8?

Not strictly, since other frameworks implement FP8 too, but TE is NVIDIA's reference implementation, and the scaling and kernel fusion work is what it saves you from writing.

Sources

#datacenter gpu#nvidia hopper#transformer engine#fp8#mixed precision#pytorch

Submit the job. Everything after that is ours.

Sign up in 60 seconds. Pay for the GPU minutes you actually use.

© 2026 Aquanode. All rights reserved.

All trademarks, logos and brand names are the property of their respective owners.