What is TFLOPS?
TFLOPS (tera floating-point operations per second) is a measure of how many trillion floating-point calculations a chip can perform each second at its theoretical peak. The H100's peak dense FP16 Tensor Core rating is 989 TFLOPS, meaning at best 989 trillion floating-point operations per second.
Floating-point operations are additions, multiplications and similar arithmetic on fractional numbers, and a fused multiply-add is conventionally counted as two. Written FLOPs (lowercase s), the term counts operations; written FLOPS, it is a rate per second. The peak comes from the datasheet: number of math units x operations per unit per clock x clock speed. It is a ceiling, not a measured speed, and real workloads land below it.
Why the precision matters
A TFLOPS figure means nothing without its number format. The same chip has separate ratings for FP32, for FP16 and BF16, and for FP8 on its Tensor Cores, and halving the bits roughly doubles the rate. The H100's dense ratings are 989 TFLOPS at FP16 and 1,979 TFLOPS at FP8.
| GPU | FP16/BF16 dense | FP8 dense | Memory bandwidth |
|---|---|---|---|
| A100 | 312 TFLOPS | none (no FP8 tensor cores) | 2,039GB/s |
| L40S | 362 TFLOPS | 733 TFLOPS | 864GB/s |
| H100 | 989 TFLOPS | 1,979 TFLOPS | 3.35TB/s |
| H200 | 989 TFLOPS | 1,979 TFLOPS | 4.8TB/s |
| MI300X | 1,307.4 TFLOPS | 2,614.9 TFLOPS | 5.3TB/s |
| B200 | 2,250 TFLOPS | 4,500 TFLOPS | 8TB/s |
Vendors also publish a second, larger figure "with sparsity", which assumes a 2:4 structured-sparsity pattern in the weights and is about double the dense rating. NVIDIA's H100 page lists 1,979 TFLOPS for FP16 and 3,958 for FP8 with sparsity, and halving them gives the dense 989 and 1,979 in the table. The same 1,979 appears as both the sparse FP16 figure and the dense FP8 figure, so check the precision and the sparsity footnote before comparing two cards.
Why TFLOPS alone does not predict LLM speed
Peak TFLOPS assumes the math units never wait for data. LLM decoding does the opposite. At batch size 1, each parameter is read from memory once per token (2 bytes at FP16) and used for about 2 operations, so the arithmetic intensity is about 1 operation per byte.
The H100's ratio of compute to bandwidth is 989 TFLOPS / 3.35TB/s, about 295 operations per byte (computed). The roofline model calls that the ridge point. A workload at 1 operation per byte sits far below it, using roughly 1/295, or 0.3%, of the compute and running at the speed of memory bandwidth instead.
So two GPUs with identical TFLOPS can decode at different speeds. The H100 and H200 both have 989 FP16 TFLOPS, but the H200 has 4.8TB/s of bandwidth against 3.35TB/s, which is 43% more. TFLOPS matters most when the work is arithmetic-heavy: training, fine-tuning, processing long prompts, and large-batch serving where each weight read is reused across many requests.
What it means when you pick a GPU
Use the right number for the job:
- Match the precision. Compare FP16 with FP16 and FP8 with FP8, and check that your card has FP8 tensor cores at all. Some workstation cards are published only with sparse or AI TOPS figures that cannot be compared directly with a dense TFLOPS rating.
- Training and fine-tuning are TFLOPS jobs. A common approximation is 6 FLOPs per parameter per training token (see "Scaling Laws for Neural Language Models", Kaplan et al., 2020). Training an 8B model on 1 billion tokens is 6 x 8 x 10^9 x 10^9 = 4.8 x 10^19 FLOPs. At the H100's 989 TFLOPS of BF16, even a perfect run takes 4.8 x 10^19 / 9.89 x 10^14, about 48,500 seconds or 13.5 hours. Real runs reach only a fraction of peak, and the weights, gradients and optimizer state must also fit in memory, so budget more time and more GPUs.
- Serving a model one request at a time is a bandwidth job. Look at memory bandwidth first. For the same compute, the H200 wins here, and the H100 usually has the lower hourly rate when the job is limited by arithmetic.
The GPU recommender shows which cards fit a given model, and Aquanode rents GPUs by the hour, with current rates on the pricing page.
Building on GPUs? Aquanode runs the workload.
Deploy on H100, H200, B200, A100 and MI300X across a multi-provider marketplace, without racking your own hardware or committing to one cloud's spec sheet.
See also
Roofline Model
The roofline model plots a kernel's arithmetic intensity against two hardware ceilings, memory bandwidth and arithmetic bandwidth, to show at a glance whether it's compute-bound or memory-bound. Where it came from and why GPUs need it.
Arithmetic Intensity
Arithmetic intensity is the ratio of compute operations to bytes moved in a kernel. Why it decides whether a workload is compute-bound or memory-bound, and how tricks like recomputation trade memory traffic for extra FLOPs.
Tensor Core
A Tensor Core is the GPU hardware unit that executes an entire matrix multiply-accumulate as one instruction instead of one scalar multiply at a time. How that trade unlocks NVIDIA's highest FLOP counts, and why an H100 has only four of them per SM.
HBM (High Bandwidth Memory)
HBM is stacked DRAM packaged beside a GPU die, giving data center cards several TB/s of memory bandwidth. Why LLM inference depends on it, and HBM vs GDDR.
BF16 (bfloat16)
BF16 is a 16-bit float with FP32's 8 exponent bits but only 7 mantissa bits. It is the default for training and costs 2 bytes per model parameter.
FP4 (NVFP4 and MXFP4)
FP4 is a 4-bit floating-point format that stores model weights at 0.5 bytes each. NVFP4 and MXFP4 add block scaling, and native FP4 math needs Blackwell GPUs.