The H100 is the better GPU in almost every measurable way: on NVIDIA's datasheets it has roughly 3.2 times the Tensor Core throughput of an A100 at matching precisions (sparse to sparse), about 1.6 times the memory bandwidth of the 80 GB A100, FP8 support the A100 lacks, and 900 GB/s NVLink against 600 GB/s. The A100 still makes sense when your model fits in 40 or 80 GB, runs in BF16 or FP16, and the lower price per hour matters more than speed.
TL;DR
- Architecture: A100 is Ampere (TSMC 7nm, 54.2 billion transistors). H100 is Hopper (TSMC 4N, 80 billion transistors).
- Compute: NVIDIA's H100 SXM figures are about 3.2x the A100 SXM figures for TF32, BF16, FP16 and INT8 (sparse vs sparse), and about 3.4x for FP64.
- New in Hopper: FP8 Tensor Cores and the Transformer Engine, which the A100 does not have.
- Memory: both top out at 80 GB, but H100 SXM moves it at 3.35 TB/s against 2.04 TB/s for the A100 80GB SXM.
- Power: up to 700 W for H100 SXM against 400 W for A100 SXM.
- Verdict: choose H100 for LLM training and serving, FP8 workloads and multi-GPU scaling. Choose A100 for smaller models, fine-tuning and experiments where it already fits.
Spec comparison
The A100 column comes from NVIDIA's A100 product page (80GB SXM) and the H100 column from NVIDIA's H100 product page (SXM). Tensor Core figures are with sparsity, as NVIDIA prints them; dense is half.
| Spec | A100 80GB SXM | H100 SXM |
|---|---|---|
| Architecture | Ampere | Hopper |
| Process | TSMC 7nm | TSMC 4N (customized) |
| Transistors | 54.2 billion | 80 billion |
| SMs | 108 | 132 |
| Memory | 80 GB HBM2e | 80 GB |
| Memory bandwidth | 2,039 GB/s | 3.35 TB/s |
| TF32 Tensor Core | 312 TFLOPS | 989 TFLOPS |
| BF16 Tensor Core | 624 TFLOPS | 1,979 TFLOPS |
| FP16 Tensor Core | 624 TFLOPS | 1,979 TFLOPS |
| FP8 Tensor Core | not supported | 3,958 TFLOPS |
| INT8 Tensor Core | 1,248 TOPS | 3,958 TOPS |
| FP64 Tensor Core | 19.5 TFLOPS (PCIe listing) | 67 TFLOPS |
| NVLink | 600 GB/s | 900 GB/s |
| PCIe | Gen4, 64 GB/s | Gen5, 128 GB/s |
| MIG | up to 7 instances at 10 GB | up to 7 instances at 10 GB |
| Max power | 400 W (up to 500 W in some HGX designs) | up to 700 W (configurable) |
Where the numbers come from, and where to be careful:
- The SM count, transistor count and process come from NVIDIA's Hopper architecture blog, which compares the A100 (SXM4) with the H100 (SXM5). That blog labels its H100 numbers preliminary and lists 3,000 GB/s of memory bandwidth, where the final H100 product page says 3.35 TB/s, so we use the product page for bandwidth and compute.
- The original A100 shipped with 40 GB of HBM2 at 1,555 GB/s according to the same blog. The 80 GB version is the one NVIDIA's A100 page lists today.
- The A100 page does not give an FP64 Tensor Core figure for the SXM version separately, so we show the PCIe listing; the Hopper blog shows 19.5 TFLOPS for the SXM4 A100 as well.
- Our ratios (3.2x for most precisions, 1.6x for memory bandwidth) are arithmetic on those datasheet numbers. They are peak figures, not measured speedups.
For the live side-by-side on this site, see A100 vs H100, and the GPU pages /gpu/nvidia-a100 and /gpu/nvidia-h100.
What changed from Ampere to Hopper
FP8 and the Transformer Engine
The biggest architectural change is the Transformer Engine. NVIDIA describes it as software plus Hopper Tensor Core hardware that picks between FP8 and 16-bit precision for each layer, handles the conversion and scaling between them, and uses tensor statistics to keep values inside FP8's narrow range. The A100 has no FP8 path, so a workload that benefits from FP8 can only get that benefit on Hopper or later. Our explainer, Transformer Engine and FP8, covers the mechanism, and the glossary has FP8 and Tensor Core.
If your stack runs in BF16 and never touches FP8, you still get the roughly 3.2x higher peak BF16 figure, but you should expect real gains to be lower than the peak ratio because memory bandwidth, communication and software overhead do not scale at the same rate.
Memory bandwidth, not capacity
Both parts are 80 GB class. Hopper's gain is bandwidth: 3.35 TB/s against 2,039 GB/s, about 1.6x by our arithmetic. For LLM decoding, which is largely limited by how fast weights and the KV cache stream out of memory, bandwidth often matters more than peak FLOPS. See KV cache. If you need more memory than 80 GB, neither card helps; look at the H200 (141 GB) described in H100 vs H200.
Interconnect
NVLink goes from 12 third-generation links and 600 GB/s on A100 to 18 fourth-generation links and 900 GB/s on H100, according to the Hopper blog and NVIDIA's product pages. PCIe moves from Gen4 to Gen5. For training across eight GPUs in one server, that makes tensor parallelism more efficient. Background in what is NVLink and NVLink vs PCIe.
Power and cooling
H100 SXM runs at up to 700 W, against 400 W for the A100 SXM. Per-GPU performance per watt still improves, but a dense Hopper server needs more cooling and power per rack than the Ampere server it replaces. Check your facility limits before swapping one for the other in place.
Performance: NVIDIA's published claims
Every figure below is NVIDIA's own, and NVIDIA labels most as projected. We have not reproduced them.
- Training, GPT-3 175B: "up to 4X faster training over the prior generation." Baseline: an A100 cluster with HDR InfiniBand against an H100 cluster with NDR InfiniBand. (H100 product page.)
- Inference on a 530B chatbot: "up to 30X" higher AI inference performance on the largest models, with an A100 HDR cluster as the baseline against H100 with NVLink Switch System and NDR InfiniBand. The workload is a Megatron chatbot with 128 input and 20 output tokens. (H100 product page.) A 530B model is far larger than anything that runs on one GPU, so this number measures the whole cluster, not one chip.
- Hopper architecture blog claims versus A100: up to 9x faster AI training and up to 30x faster inference on large language models, 3x for FP64 and FP32 chip to chip, and "approximately 6x" overall peak compute, which NVIDIA calls a combined estimate from SM count, per-SM speedup, FP8 and clock frequency.
- Llama 2 70B on H100 NVL: up to 5x over A100 systems (H100 NVL, the PCIe variant, not SXM).
- MLPerf Inference, September 2022: in NVIDIA's account of H100's first MLPerf round, Hopper delivered "up to 4.5x more performance than NVIDIA Ampere architecture GPUs" and credited the Transformer Engine in part for results on BERT. This was an audited submission, though NVIDIA chose the framing.
The pattern: audited per-chip results land at several times, while cluster-scale claims on giant models reach 30x because they combine faster chips with faster networking and bigger NVLink domains. Use the former for single-GPU decisions.
When the A100 is still the right call
Choose A100 when:
- Your model and activations fit in 40 or 80 GB and you train or serve in BF16 or FP16.
- You fine-tune with parameter-efficient methods (LoRA) on small to mid-sized models, where the job is short and speed matters less than the hourly price.
- You use MIG to split one GPU into up to seven instances for many small inference jobs. Both GPUs support MIG, so this is not an H100 advantage, and the cheaper card can be the better fit.
- Your code base is tuned for Ampere and you have no FP8 plan.
- You need to match an existing A100 fleet's behavior for reproducibility.
Choose H100 when:
- You train or serve LLMs where FP8 reduces memory and increases throughput.
- Your parallelism needs the 900 GB/s NVLink and Gen5 PCIe.
- Time to result, not cost per hour, is the constraint.
- You want the broadest software support for current frameworks and libraries.
If you are weighing H100 against newer chips, see H200 vs B200 vs GB200, and the full chip map in the datacenter GPU guide. For the H100 variants (SXM, NVL, PCIe), read H100 and H200 SXM vs NVL vs PCIe.
Cost: compare per job, not per hour
An H100 costs more per hour than an A100, but its higher throughput can mean a lower cost per token or per training run. The comparison depends on how much of the peak you actually reach, so measure it: run your own workload for a short fixed time on each, record tokens per second or steps per second, and divide the live hourly price by your measured throughput. NVIDIA's data suggests a peak ratio of roughly 3x for BF16 and more with FP8, so if the H100 hourly price is under about three times the A100 price and you are compute bound, H100 should come out ahead per job. If you are bound by something else (CPU data loading, small batch sizes, network), the gap narrows and the A100 can win. We do not type prices here; use the live box below for the H100 and H200 rates.
Rent today
Aquanode manages and optimizes GPUs for training and inference workloads, and you can rent the GPUs below on demand. The box shows live availability; "None right now" means there is no offer at the moment.
What's next
After Hopper came Blackwell. NVIDIA's DGX B200 page claims 3x the training and 15x the inference performance of DGX H100 (vendor projection). Read H200 vs B200 vs GB200 for how those claims hold up against audited MLPerf results, and Rubin vs Blackwell vs Hopper for the roadmap.
FAQ
How much faster is H100 than A100?
On NVIDIA's datasheets, about 3.2x at matching Tensor Core precisions (sparse vs sparse) and about 1.6x on memory bandwidth. NVIDIA's claims for specific workloads range from 4.5x in its first MLPerf inference round to 9x for LLM training on large clusters. Your result depends on whether you are compute bound.
Does A100 support FP8?
No. FP8 Tensor Cores arrived with Hopper. NVIDIA's A100 page lists TF32, BF16, FP16 and INT8 as the Tensor Core precisions.
Is A100 still worth using in 2026?
Yes for models that fit in its memory, for BF16 or FP16 work and for cost-sensitive fine-tuning. For FP8 serving or large-scale training, H100 and newer parts are the better fit.
Do A100 and H100 both have 80 GB?
Both are available with 80 GB. The original A100 launched with 40 GB, per NVIDIA's Hopper blog, and the H100 NVL variant has 94 GB.
What is the difference in power draw?
NVIDIA lists up to 700 W for H100 SXM and 400 W for the A100 SXM, with some HGX A100 designs supporting up to 500 W.
Sources
- NVIDIA A100 product page: https://www.nvidia.com/en-us/data-center/a100/
- NVIDIA H100 product page: https://www.nvidia.com/en-us/data-center/h100/
- NVIDIA technical blog, Hopper architecture in depth: https://developer.nvidia.com/blog/nvidia-hopper-architecture-in-depth/
- NVIDIA blog, Hopper MLPerf inference debut (September 8, 2022): https://blogs.nvidia.com/blog/hopper-mlperf-inference/
- NVIDIA DGX B200 product page: https://www.nvidia.com/en-us/data-center/dgx-b200/