What is TFLOPS?

TFLOPS (tera floating-point operations per second) is a measure of how many trillion floating-point calculations a chip can perform each second at its theoretical peak. The H100's peak dense FP16 Tensor Core rating is 989 TFLOPS, meaning at best 989 trillion floating-point operations per second.

Floating-point operations are additions, multiplications and similar arithmetic on fractional numbers, and a fused multiply-add is conventionally counted as two. Written FLOPs (lowercase s), the term counts operations; written FLOPS, it is a rate per second. The peak comes from the datasheet: number of math units x operations per unit per clock x clock speed. It is a ceiling, not a measured speed, and real workloads land below it.

Why the precision matters

A TFLOPS figure means nothing without its number format. The same chip has separate ratings for FP32, for FP16 and BF16, and for FP8 on its Tensor Cores, and halving the bits roughly doubles the rate. The H100's dense ratings are 989 TFLOPS at FP16 and 1,979 TFLOPS at FP8.

GPUFP16/BF16 denseFP8 denseMemory bandwidth
A100312 TFLOPSnone (no FP8 tensor cores)2,039GB/s
L40S362 TFLOPS733 TFLOPS864GB/s
H100989 TFLOPS1,979 TFLOPS3.35TB/s
H200989 TFLOPS1,979 TFLOPS4.8TB/s
MI300X1,307.4 TFLOPS2,614.9 TFLOPS5.3TB/s
B2002,250 TFLOPS4,500 TFLOPS8TB/s

Vendors also publish a second, larger figure "with sparsity", which assumes a 2:4 structured-sparsity pattern in the weights and is about double the dense rating. NVIDIA's H100 page lists 1,979 TFLOPS for FP16 and 3,958 for FP8 with sparsity, and halving them gives the dense 989 and 1,979 in the table. The same 1,979 appears as both the sparse FP16 figure and the dense FP8 figure, so check the precision and the sparsity footnote before comparing two cards.

Why TFLOPS alone does not predict LLM speed

Peak TFLOPS assumes the math units never wait for data. LLM decoding does the opposite. At batch size 1, each parameter is read from memory once per token (2 bytes at FP16) and used for about 2 operations, so the arithmetic intensity is about 1 operation per byte.

The H100's ratio of compute to bandwidth is 989 TFLOPS / 3.35TB/s, about 295 operations per byte (computed). The roofline model calls that the ridge point. A workload at 1 operation per byte sits far below it, using roughly 1/295, or 0.3%, of the compute and running at the speed of memory bandwidth instead.

So two GPUs with identical TFLOPS can decode at different speeds. The H100 and H200 both have 989 FP16 TFLOPS, but the H200 has 4.8TB/s of bandwidth against 3.35TB/s, which is 43% more. TFLOPS matters most when the work is arithmetic-heavy: training, fine-tuning, processing long prompts, and large-batch serving where each weight read is reused across many requests.

What it means when you pick a GPU

Use the right number for the job:

  • Match the precision. Compare FP16 with FP16 and FP8 with FP8, and check that your card has FP8 tensor cores at all. Some workstation cards are published only with sparse or AI TOPS figures that cannot be compared directly with a dense TFLOPS rating.
  • Training and fine-tuning are TFLOPS jobs. A common approximation is 6 FLOPs per parameter per training token (see "Scaling Laws for Neural Language Models", Kaplan et al., 2020). Training an 8B model on 1 billion tokens is 6 x 8 x 10^9 x 10^9 = 4.8 x 10^19 FLOPs. At the H100's 989 TFLOPS of BF16, even a perfect run takes 4.8 x 10^19 / 9.89 x 10^14, about 48,500 seconds or 13.5 hours. Real runs reach only a fraction of peak, and the weights, gradients and optimizer state must also fit in memory, so budget more time and more GPUs.
  • Serving a model one request at a time is a bandwidth job. Look at memory bandwidth first. For the same compute, the H200 wins here, and the H100 usually has the lower hourly rate when the job is limited by arithmetic.

The GPU recommender shows which cards fit a given model, and Aquanode rents GPUs by the hour, with current rates on the pricing page.

Building on GPUs? Aquanode runs the workload.

Deploy on H100, H200, B200, A100 and MI300X across a multi-provider marketplace, without racking your own hardware or committing to one cloud's spec sheet.

See also

Submit the job. Everything after that is ours.

Sign up in 60 seconds. Pay for the GPU minutes you actually use.

© 2026 Aquanode. All rights reserved.

All trademarks, logos and brand names are the property of their respective owners.