H100 vs H200: Specs, Benchmarks and Which to Pick (2026)

Back
Team Aquanode

Team Aquanode

Sarthak Vaish

Updated OCTOBER 8, 2026Published OCTOBER 8, 2026

The H200 is an H100 with a bigger, faster memory system: 141 GB of HBM3e at 4.8 TB/s against the H100 SXM's 80 GB of HBM3 at 3.35 TB/s, with the same tensor-core compute and the same 700 W power class. Pick the H200 when memory capacity or bandwidth limits you, and the H100 when your model already fits and the job is compute-bound.

This comparison covers:

  • A spec table taken from NVIDIA's pages
  • Why identical compute but different memory changes results
  • NVIDIA's MLPerf and product-page deltas, with their conditions
  • Which GPU wins for inference, training, fine-tuning and long context

TL;DR

  • Same compute. NVIDIA's tables list identical tensor-core TFLOPS for the H100 SXM and H200 SXM (for example 3,958 TFLOPS FP8 with sparsity).
  • Different memory. H200: 141 GB HBM3e, 4.8 TB/s. H100: 80 GB HBM3, 3.35 TB/s. That is about 1.76x the capacity and 1.43x the bandwidth (our arithmetic).
  • Measured gain. NVIDIA reports up to 28% better Llama 2 70B performance for the H200 at the same 700 W in MLPerf Inference v4.0, and larger gains on its product page where the H200 runs larger batches.
  • Verdict: the H200 for 70B-class inference, long contexts and big batches; the H100 for compute-bound work and models that fit in 80 GB, especially when its hourly price is meaningfully lower.

Deep dives: the H200 guide, the form factor guide, and the datacenter GPU overview. For a live side-by-side with current prices, use the H100 vs H200 comparison page.

H100 vs H200 spec table

SXM versions, from NVIDIA's H100 and H200 product pages. Tensor-core figures are quoted by NVIDIA with sparsity; dense is half (our arithmetic).

SpecH100 SXMH200 SXMDifference
GPU memory80 GB HBM3141 GB HBM3eAbout 1.76x
Memory bandwidth3.35 TB/s4.8 TB/sAbout 1.43x
FP8 tensor (sparsity)3,958 TFLOPS3,958 TFLOPSSame
BF16 tensor (sparsity)1,979 TFLOPS1,979 TFLOPSSame
TF32 tensor (sparsity)989 TFLOPS989 TFLOPSSame
FP64 / FP64 tensor34 / 67 TFLOPS34 / 67 TFLOPSSame
Max powerUp to 700 W (configurable)Up to 700 W (configurable)Same
NVLink900 GB/s900 GB/sSame
PCIeGen 5, 128 GB/sGen 5, 128 GB/sSame
MIGUp to 7 at 10 GBUp to 7 at 18 GBLarger slices
Server optionsHGX H100HGX H200, 4 or 8 GPUsHGX boards compatible

The pattern is clear: every row that describes compute, power or interconnect is unchanged, and the two memory rows are what moved. NVIDIA's H200 press release states that HGX H200 boards are compatible with both the hardware and software of HGX H100 systems.

The PCIe-format versions differ more. The H100 NVL has 94 GB at 3.9 TB/s and the H200 NVL has 141 GB at 4.8 TB/s; NVIDIA claims up to 1.7x faster LLM inference for the H200 NVL over the H100 NVL, without listing test conditions. Details are in SXM vs NVL vs PCIe.

Why identical compute still gives different speed

LLM inference has two phases:

  • Prefill processes the prompt in parallel. It is compute-bound, so a chip with the same tensor-core throughput performs the same.
  • Decode generates one token at a time. For each token the GPU reads the model weights and the KV cache from memory, so speed is set by memory bandwidth. More bandwidth means faster decode, at the same compute.

Capacity matters too. If a model plus its KV cache does not fit in 80 GB, you must split it across GPUs with tensor or pipeline parallelism, which adds communication on every layer (see tensor parallelism). If it fits on one H200, you avoid that. NVIDIA's MLPerf write-up gives this as a reason for the H200's Llama 2 70B result: the extra memory removes the need for tensor or pipeline parallelism, and the higher bandwidth relieves memory-bound bottlenecks.

Training sits differently. Large-scale training is mostly compute and interconnect bound, so the H200's advantage is smaller there, unless memory pressure forces you into more aggressive sharding or smaller micro-batches on the H100.

Benchmarks: what NVIDIA has published

All numbers are NVIDIA's, with their stated conditions.

MLPerf Inference v4.0, Llama 2 70B (NVIDIA blog, results retrieved March 27, 2024):

ConfigurationH200 vs H100
H200 at the same 700 W TDPUp to 28% better
H200 at 1,000 W, custom thermal designUp to 45% better (43% server, 45% offline)

The 1,000 W configuration is not a standard HGX setup, so the 28% figure is the one to plan around. NVIDIA's footnote lists the H200 submissions as 4.0-0062 and 4.0-0068.

H200 product page claims (vendor's claims):

  • Llama2 70B: 1.9x faster than H100 SXM, with input length 2K, output length 128, and batch size 32 on the H200 against 8 on the H100 per GPU.
  • GPT-3 175B: 1.6x faster, 8x H200 against 8x H100, input length 80, output length 200, batch size 128 against 64.
  • Up to 2x faster LLM inference than H100 for models such as Llama2.

Look at the batch sizes: 32 against 8, and 128 against 64. These compare each GPU at the batch size its memory permits, which is realistic for production serving but not an equal-batch test. The H200's larger memory lets you run larger batches, so a deployment that can use them captures the larger gain, while a latency-bound deployment at small batch will see closer to the memory-bandwidth ratio of 1.43x or less.

We have not run our own H100 vs H200 benchmark for this post, and we do not quote third-party benchmarks here. Run your own model at your real batch size and context length.

Which GPU wins for each workload

WorkloadBetter pickWhy
70B model in FP8, one GPUH200About 70 GB of weights (our arithmetic) leaves ample room for KV cache in 141 GB; on 80 GB it barely fits
70B model in BF16H200 (2 GPUs) or H100 (2 GPUs)About 140 GB of weights needs two GPUs on either, but the H200 pair has far more cache room
Long-context servingH200KV cache grows with context; more VRAM and bandwidth
Small models (8B and under)H100The model fits easily, compute is enough, cost matters more
Prefill-heavy or batch scoringH100Compute-bound, identical tensor-core throughput
Pre-trainingH100 or H200Same compute; H200 only helps if memory forces sharding
Fine-tuning 30B to 70BH200More room for weights, optimizer state and activations
Mixed workloads on MIG slicesH20018 GB slices against 10 GB

For the FP8 point, see Transformer Engine and FP8.

Power and infrastructure

There is nothing to re-engineer. Both SXM parts are configurable up to 700 W, both use 900 GB/s NVLink and PCIe Gen 5, and NVIDIA says HGX H200 boards are compatible with HGX H100 hardware and software. A cluster designed for H100 can in principle take H200 baseboards, subject to your system vendor's confirmation. For how NVSwitch connects the GPUs, read what is NVLink.

Cost: how to decide

Price and speed together decide this, and the price moves, so the live box below shows today's from-price per GPU-hour for each. To compare:

  1. Measure tokens per second on each GPU with your model at your batch size and context length.
  2. Multiply by 3,600 for tokens per GPU-hour.
  3. Divide each GPU's live hourly price by its tokens per GPU-hour.

The H200 is the better value when its throughput gain exceeds its price premium over the H100. If your test shows a gain below the price ratio, the H100 is cheaper per token. If the model needs two H100s to fit but only one H200, compare per model replica, not per GPU.

Rent today

Aquanode manages and optimizes GPUs for training and inference workloads, and you can rent the GPUs in the box below on demand.

Specs and availability: H100, H200, and /pricing. For the whole market, see the GPU index.

What's next

If neither fits, Blackwell raises memory and FP4 throughput: see the B200 guide and H200 vs B200 vs GB200.

FAQ

Is the H200 faster than the H100?

Only where memory matters. Tensor-core throughput is identical in NVIDIA's tables. NVIDIA reports up to 28% better Llama 2 70B performance at the same 700 W in MLPerf Inference v4.0, and up to 1.9x on its product page with larger batches.

How much more memory does the H200 have?

141 GB of HBM3e against 80 GB of HBM3 on the H100 SXM, about 1.76x, with 4.8 TB/s against 3.35 TB/s of bandwidth, about 1.43x.

Should I upgrade from H100 to H200?

If your model, KV cache or batch size is limited by 80 GB, yes. If your job is compute-bound or already fits comfortably, the gain is small and the H100 may be cheaper per token.

Do H100 and H200 use the same servers?

NVIDIA says HGX H200 boards are compatible with the hardware and software of HGX H100 systems. Both SXM parts run at up to 700 W.

Is H200 better for training?

Not by much on compute, since the tensor-core numbers match. It can help when memory limits your sharding or micro-batch size.

What about the PCIe versions?

The H100 NVL has 94 GB and the H200 NVL has 141 GB. NVIDIA claims up to 1.7x faster LLM inference for the H200 NVL; see the form factor guide.

Sources

#datacenter gpu#nvidia hopper#h100 vs h200#gpu comparison#llm inference#hbm3e

Submit the job. Everything after that is ours.

Sign up in 60 seconds. Pay for the GPU minutes you actually use.

© 2026 Aquanode. All rights reserved.

All trademarks, logos and brand names are the property of their respective owners.