A100 vs H100: Specs, Benchmarks, Which to Pick (2026)

Back
Team Aquanode

Team Aquanode

Sarthak Vaish

Updated OCTOBER 8, 2026Published OCTOBER 8, 2026

The H100 is the better GPU in almost every measurable way: on NVIDIA's datasheets it has roughly 3.2 times the Tensor Core throughput of an A100 at matching precisions (sparse to sparse), about 1.6 times the memory bandwidth of the 80 GB A100, FP8 support the A100 lacks, and 900 GB/s NVLink against 600 GB/s. The A100 still makes sense when your model fits in 40 or 80 GB, runs in BF16 or FP16, and the lower price per hour matters more than speed.

TL;DR

  • Architecture: A100 is Ampere (TSMC 7nm, 54.2 billion transistors). H100 is Hopper (TSMC 4N, 80 billion transistors).
  • Compute: NVIDIA's H100 SXM figures are about 3.2x the A100 SXM figures for TF32, BF16, FP16 and INT8 (sparse vs sparse), and about 3.4x for FP64.
  • New in Hopper: FP8 Tensor Cores and the Transformer Engine, which the A100 does not have.
  • Memory: both top out at 80 GB, but H100 SXM moves it at 3.35 TB/s against 2.04 TB/s for the A100 80GB SXM.
  • Power: up to 700 W for H100 SXM against 400 W for A100 SXM.
  • Verdict: choose H100 for LLM training and serving, FP8 workloads and multi-GPU scaling. Choose A100 for smaller models, fine-tuning and experiments where it already fits.

Spec comparison

The A100 column comes from NVIDIA's A100 product page (80GB SXM) and the H100 column from NVIDIA's H100 product page (SXM). Tensor Core figures are with sparsity, as NVIDIA prints them; dense is half.

SpecA100 80GB SXMH100 SXM
ArchitectureAmpereHopper
ProcessTSMC 7nmTSMC 4N (customized)
Transistors54.2 billion80 billion
SMs108132
Memory80 GB HBM2e80 GB
Memory bandwidth2,039 GB/s3.35 TB/s
TF32 Tensor Core312 TFLOPS989 TFLOPS
BF16 Tensor Core624 TFLOPS1,979 TFLOPS
FP16 Tensor Core624 TFLOPS1,979 TFLOPS
FP8 Tensor Corenot supported3,958 TFLOPS
INT8 Tensor Core1,248 TOPS3,958 TOPS
FP64 Tensor Core19.5 TFLOPS (PCIe listing)67 TFLOPS
NVLink600 GB/s900 GB/s
PCIeGen4, 64 GB/sGen5, 128 GB/s
MIGup to 7 instances at 10 GBup to 7 instances at 10 GB
Max power400 W (up to 500 W in some HGX designs)up to 700 W (configurable)

Where the numbers come from, and where to be careful:

  • The SM count, transistor count and process come from NVIDIA's Hopper architecture blog, which compares the A100 (SXM4) with the H100 (SXM5). That blog labels its H100 numbers preliminary and lists 3,000 GB/s of memory bandwidth, where the final H100 product page says 3.35 TB/s, so we use the product page for bandwidth and compute.
  • The original A100 shipped with 40 GB of HBM2 at 1,555 GB/s according to the same blog. The 80 GB version is the one NVIDIA's A100 page lists today.
  • The A100 page does not give an FP64 Tensor Core figure for the SXM version separately, so we show the PCIe listing; the Hopper blog shows 19.5 TFLOPS for the SXM4 A100 as well.
  • Our ratios (3.2x for most precisions, 1.6x for memory bandwidth) are arithmetic on those datasheet numbers. They are peak figures, not measured speedups.

For the live side-by-side on this site, see A100 vs H100, and the GPU pages /gpu/nvidia-a100 and /gpu/nvidia-h100.

What changed from Ampere to Hopper

FP8 and the Transformer Engine

The biggest architectural change is the Transformer Engine. NVIDIA describes it as software plus Hopper Tensor Core hardware that picks between FP8 and 16-bit precision for each layer, handles the conversion and scaling between them, and uses tensor statistics to keep values inside FP8's narrow range. The A100 has no FP8 path, so a workload that benefits from FP8 can only get that benefit on Hopper or later. Our explainer, Transformer Engine and FP8, covers the mechanism, and the glossary has FP8 and Tensor Core.

If your stack runs in BF16 and never touches FP8, you still get the roughly 3.2x higher peak BF16 figure, but you should expect real gains to be lower than the peak ratio because memory bandwidth, communication and software overhead do not scale at the same rate.

Memory bandwidth, not capacity

Both parts are 80 GB class. Hopper's gain is bandwidth: 3.35 TB/s against 2,039 GB/s, about 1.6x by our arithmetic. For LLM decoding, which is largely limited by how fast weights and the KV cache stream out of memory, bandwidth often matters more than peak FLOPS. See KV cache. If you need more memory than 80 GB, neither card helps; look at the H200 (141 GB) described in H100 vs H200.

Interconnect

NVLink goes from 12 third-generation links and 600 GB/s on A100 to 18 fourth-generation links and 900 GB/s on H100, according to the Hopper blog and NVIDIA's product pages. PCIe moves from Gen4 to Gen5. For training across eight GPUs in one server, that makes tensor parallelism more efficient. Background in what is NVLink and NVLink vs PCIe.

Power and cooling

H100 SXM runs at up to 700 W, against 400 W for the A100 SXM. Per-GPU performance per watt still improves, but a dense Hopper server needs more cooling and power per rack than the Ampere server it replaces. Check your facility limits before swapping one for the other in place.

Performance: NVIDIA's published claims

Every figure below is NVIDIA's own, and NVIDIA labels most as projected. We have not reproduced them.

  • Training, GPT-3 175B: "up to 4X faster training over the prior generation." Baseline: an A100 cluster with HDR InfiniBand against an H100 cluster with NDR InfiniBand. (H100 product page.)
  • Inference on a 530B chatbot: "up to 30X" higher AI inference performance on the largest models, with an A100 HDR cluster as the baseline against H100 with NVLink Switch System and NDR InfiniBand. The workload is a Megatron chatbot with 128 input and 20 output tokens. (H100 product page.) A 530B model is far larger than anything that runs on one GPU, so this number measures the whole cluster, not one chip.
  • Hopper architecture blog claims versus A100: up to 9x faster AI training and up to 30x faster inference on large language models, 3x for FP64 and FP32 chip to chip, and "approximately 6x" overall peak compute, which NVIDIA calls a combined estimate from SM count, per-SM speedup, FP8 and clock frequency.
  • Llama 2 70B on H100 NVL: up to 5x over A100 systems (H100 NVL, the PCIe variant, not SXM).
  • MLPerf Inference, September 2022: in NVIDIA's account of H100's first MLPerf round, Hopper delivered "up to 4.5x more performance than NVIDIA Ampere architecture GPUs" and credited the Transformer Engine in part for results on BERT. This was an audited submission, though NVIDIA chose the framing.

The pattern: audited per-chip results land at several times, while cluster-scale claims on giant models reach 30x because they combine faster chips with faster networking and bigger NVLink domains. Use the former for single-GPU decisions.

When the A100 is still the right call

Choose A100 when:

  • Your model and activations fit in 40 or 80 GB and you train or serve in BF16 or FP16.
  • You fine-tune with parameter-efficient methods (LoRA) on small to mid-sized models, where the job is short and speed matters less than the hourly price.
  • You use MIG to split one GPU into up to seven instances for many small inference jobs. Both GPUs support MIG, so this is not an H100 advantage, and the cheaper card can be the better fit.
  • Your code base is tuned for Ampere and you have no FP8 plan.
  • You need to match an existing A100 fleet's behavior for reproducibility.

Choose H100 when:

  • You train or serve LLMs where FP8 reduces memory and increases throughput.
  • Your parallelism needs the 900 GB/s NVLink and Gen5 PCIe.
  • Time to result, not cost per hour, is the constraint.
  • You want the broadest software support for current frameworks and libraries.

If you are weighing H100 against newer chips, see H200 vs B200 vs GB200, and the full chip map in the datacenter GPU guide. For the H100 variants (SXM, NVL, PCIe), read H100 and H200 SXM vs NVL vs PCIe.

Cost: compare per job, not per hour

An H100 costs more per hour than an A100, but its higher throughput can mean a lower cost per token or per training run. The comparison depends on how much of the peak you actually reach, so measure it: run your own workload for a short fixed time on each, record tokens per second or steps per second, and divide the live hourly price by your measured throughput. NVIDIA's data suggests a peak ratio of roughly 3x for BF16 and more with FP8, so if the H100 hourly price is under about three times the A100 price and you are compute bound, H100 should come out ahead per job. If you are bound by something else (CPU data loading, small batch sizes, network), the gap narrows and the A100 can win. We do not type prices here; use the live box below for the H100 and H200 rates.

Rent today

Aquanode manages and optimizes GPUs for training and inference workloads, and you can rent the GPUs below on demand. The box shows live availability; "None right now" means there is no offer at the moment.

What's next

After Hopper came Blackwell. NVIDIA's DGX B200 page claims 3x the training and 15x the inference performance of DGX H100 (vendor projection). Read H200 vs B200 vs GB200 for how those claims hold up against audited MLPerf results, and Rubin vs Blackwell vs Hopper for the roadmap.

FAQ

How much faster is H100 than A100?

On NVIDIA's datasheets, about 3.2x at matching Tensor Core precisions (sparse vs sparse) and about 1.6x on memory bandwidth. NVIDIA's claims for specific workloads range from 4.5x in its first MLPerf inference round to 9x for LLM training on large clusters. Your result depends on whether you are compute bound.

Does A100 support FP8?

No. FP8 Tensor Cores arrived with Hopper. NVIDIA's A100 page lists TF32, BF16, FP16 and INT8 as the Tensor Core precisions.

Is A100 still worth using in 2026?

Yes for models that fit in its memory, for BF16 or FP16 work and for cost-sensitive fine-tuning. For FP8 serving or large-scale training, H100 and newer parts are the better fit.

Do A100 and H100 both have 80 GB?

Both are available with 80 GB. The original A100 launched with 40 GB, per NVIDIA's Hopper blog, and the H100 NVL variant has 94 GB.

What is the difference in power draw?

NVIDIA lists up to 700 W for H100 SXM and 400 W for the A100 SXM, with some HGX A100 designs supporting up to 500 W.

Sources

#datacenter gpu#nvidia hopper#nvidia ampere#a100 vs h100#h100#a100

Related reading

V100 vs A100: Specs and Is the V100 Worth It in 2026

What the NVIDIA V100 is, its specs, and how it compares with the A100 on memory, bandwidth, Tensor Cores and NVLink, plus whether it is worth using in 2026.

Best GPU for LLM Inference: A Segment-by-Segment Guide

The best GPU for LLM inference depends on your segment. H100 for production serving, A100 for training value, L40S for local dev, with real specs and pricing.

What Is the Difference Between AMD and NVIDIA GPUs for AI?

AMD vs NVIDIA GPUs for AI: architecture, CUDA vs ROCm, memory, pricing, and when each one is actually the right call for training and inference.

GPU Monitoring Tools Compared: nvidia-smi and More

A practical comparison of GPU monitoring tools: nvidia-smi, gpustat, nvtop, nvitop and jupyterlab-nvdashboard, with install commands and when to use each.

H100 and H200: SXM vs NVL vs PCIe Explained (2026)

H100 and H200 come as SXM, NVL and PCIe cards. Specs, NVLink bandwidth, power and which form factor fits inference, training or an air-cooled rack.

NVIDIA GH200 Grace Hopper Superchip: Specs and Guide (2026)

What the NVIDIA GH200 Grace Hopper Superchip is, its specs (96 GB HBM3 or 144 GB HBM3e, 480 GB LPDDR5X, NVLink-C2C), and when it beats an H100.

NVIDIA H20: Specs, Export Status and Guide (2026)

What the NVIDIA H20 is, its reported specs against the H100 and H200, and its export-license history with dates, as of October 2026.

faster-whisper vs insanely-fast-whisper vs WhisperX

faster-whisper, insanely-fast-whisper and WhisperX all run the same Whisper weights. Here is how their engines differ and which one to pick.

Submit the job. Everything after that is ours.

Sign up in 60 seconds. Pay for the GPU minutes you actually use.

© 2026 Aquanode. All rights reserved.

All trademarks, logos and brand names are the property of their respective owners.