The H200 is an H100 with a bigger, faster memory system: 141 GB of HBM3e at 4.8 TB/s against the H100 SXM's 80 GB of HBM3 at 3.35 TB/s, with the same tensor-core compute and the same 700 W power class. Pick the H200 when memory capacity or bandwidth limits you, and the H100 when your model already fits and the job is compute-bound.
This comparison covers:
- A spec table taken from NVIDIA's pages
- Why identical compute but different memory changes results
- NVIDIA's MLPerf and product-page deltas, with their conditions
- Which GPU wins for inference, training, fine-tuning and long context
TL;DR
- Same compute. NVIDIA's tables list identical tensor-core TFLOPS for the H100 SXM and H200 SXM (for example 3,958 TFLOPS FP8 with sparsity).
- Different memory. H200: 141 GB HBM3e, 4.8 TB/s. H100: 80 GB HBM3, 3.35 TB/s. That is about 1.76x the capacity and 1.43x the bandwidth (our arithmetic).
- Measured gain. NVIDIA reports up to 28% better Llama 2 70B performance for the H200 at the same 700 W in MLPerf Inference v4.0, and larger gains on its product page where the H200 runs larger batches.
- Verdict: the H200 for 70B-class inference, long contexts and big batches; the H100 for compute-bound work and models that fit in 80 GB, especially when its hourly price is meaningfully lower.
Deep dives: the H200 guide, the form factor guide, and the datacenter GPU overview. For a live side-by-side with current prices, use the H100 vs H200 comparison page.
H100 vs H200 spec table
SXM versions, from NVIDIA's H100 and H200 product pages. Tensor-core figures are quoted by NVIDIA with sparsity; dense is half (our arithmetic).
| Spec | H100 SXM | H200 SXM | Difference |
|---|---|---|---|
| GPU memory | 80 GB HBM3 | 141 GB HBM3e | About 1.76x |
| Memory bandwidth | 3.35 TB/s | 4.8 TB/s | About 1.43x |
| FP8 tensor (sparsity) | 3,958 TFLOPS | 3,958 TFLOPS | Same |
| BF16 tensor (sparsity) | 1,979 TFLOPS | 1,979 TFLOPS | Same |
| TF32 tensor (sparsity) | 989 TFLOPS | 989 TFLOPS | Same |
| FP64 / FP64 tensor | 34 / 67 TFLOPS | 34 / 67 TFLOPS | Same |
| Max power | Up to 700 W (configurable) | Up to 700 W (configurable) | Same |
| NVLink | 900 GB/s | 900 GB/s | Same |
| PCIe | Gen 5, 128 GB/s | Gen 5, 128 GB/s | Same |
| MIG | Up to 7 at 10 GB | Up to 7 at 18 GB | Larger slices |
| Server options | HGX H100 | HGX H200, 4 or 8 GPUs | HGX boards compatible |
The pattern is clear: every row that describes compute, power or interconnect is unchanged, and the two memory rows are what moved. NVIDIA's H200 press release states that HGX H200 boards are compatible with both the hardware and software of HGX H100 systems.
The PCIe-format versions differ more. The H100 NVL has 94 GB at 3.9 TB/s and the H200 NVL has 141 GB at 4.8 TB/s; NVIDIA claims up to 1.7x faster LLM inference for the H200 NVL over the H100 NVL, without listing test conditions. Details are in SXM vs NVL vs PCIe.
Why identical compute still gives different speed
LLM inference has two phases:
- Prefill processes the prompt in parallel. It is compute-bound, so a chip with the same tensor-core throughput performs the same.
- Decode generates one token at a time. For each token the GPU reads the model weights and the KV cache from memory, so speed is set by memory bandwidth. More bandwidth means faster decode, at the same compute.
Capacity matters too. If a model plus its KV cache does not fit in 80 GB, you must split it across GPUs with tensor or pipeline parallelism, which adds communication on every layer (see tensor parallelism). If it fits on one H200, you avoid that. NVIDIA's MLPerf write-up gives this as a reason for the H200's Llama 2 70B result: the extra memory removes the need for tensor or pipeline parallelism, and the higher bandwidth relieves memory-bound bottlenecks.
Training sits differently. Large-scale training is mostly compute and interconnect bound, so the H200's advantage is smaller there, unless memory pressure forces you into more aggressive sharding or smaller micro-batches on the H100.
Benchmarks: what NVIDIA has published
All numbers are NVIDIA's, with their stated conditions.
MLPerf Inference v4.0, Llama 2 70B (NVIDIA blog, results retrieved March 27, 2024):
| Configuration | H200 vs H100 |
|---|---|
| H200 at the same 700 W TDP | Up to 28% better |
| H200 at 1,000 W, custom thermal design | Up to 45% better (43% server, 45% offline) |
The 1,000 W configuration is not a standard HGX setup, so the 28% figure is the one to plan around. NVIDIA's footnote lists the H200 submissions as 4.0-0062 and 4.0-0068.
H200 product page claims (vendor's claims):
- Llama2 70B: 1.9x faster than H100 SXM, with input length 2K, output length 128, and batch size 32 on the H200 against 8 on the H100 per GPU.
- GPT-3 175B: 1.6x faster, 8x H200 against 8x H100, input length 80, output length 200, batch size 128 against 64.
- Up to 2x faster LLM inference than H100 for models such as Llama2.
Look at the batch sizes: 32 against 8, and 128 against 64. These compare each GPU at the batch size its memory permits, which is realistic for production serving but not an equal-batch test. The H200's larger memory lets you run larger batches, so a deployment that can use them captures the larger gain, while a latency-bound deployment at small batch will see closer to the memory-bandwidth ratio of 1.43x or less.
We have not run our own H100 vs H200 benchmark for this post, and we do not quote third-party benchmarks here. Run your own model at your real batch size and context length.
Which GPU wins for each workload
| Workload | Better pick | Why |
|---|---|---|
| 70B model in FP8, one GPU | H200 | About 70 GB of weights (our arithmetic) leaves ample room for KV cache in 141 GB; on 80 GB it barely fits |
| 70B model in BF16 | H200 (2 GPUs) or H100 (2 GPUs) | About 140 GB of weights needs two GPUs on either, but the H200 pair has far more cache room |
| Long-context serving | H200 | KV cache grows with context; more VRAM and bandwidth |
| Small models (8B and under) | H100 | The model fits easily, compute is enough, cost matters more |
| Prefill-heavy or batch scoring | H100 | Compute-bound, identical tensor-core throughput |
| Pre-training | H100 or H200 | Same compute; H200 only helps if memory forces sharding |
| Fine-tuning 30B to 70B | H200 | More room for weights, optimizer state and activations |
| Mixed workloads on MIG slices | H200 | 18 GB slices against 10 GB |
For the FP8 point, see Transformer Engine and FP8.
Power and infrastructure
There is nothing to re-engineer. Both SXM parts are configurable up to 700 W, both use 900 GB/s NVLink and PCIe Gen 5, and NVIDIA says HGX H200 boards are compatible with HGX H100 hardware and software. A cluster designed for H100 can in principle take H200 baseboards, subject to your system vendor's confirmation. For how NVSwitch connects the GPUs, read what is NVLink.
Cost: how to decide
Price and speed together decide this, and the price moves, so the live box below shows today's from-price per GPU-hour for each. To compare:
- Measure tokens per second on each GPU with your model at your batch size and context length.
- Multiply by 3,600 for tokens per GPU-hour.
- Divide each GPU's live hourly price by its tokens per GPU-hour.
The H200 is the better value when its throughput gain exceeds its price premium over the H100. If your test shows a gain below the price ratio, the H100 is cheaper per token. If the model needs two H100s to fit but only one H200, compare per model replica, not per GPU.
Rent today
Aquanode manages and optimizes GPUs for training and inference workloads, and you can rent the GPUs in the box below on demand.
Specs and availability: H100, H200, and /pricing. For the whole market, see the GPU index.
What's next
If neither fits, Blackwell raises memory and FP4 throughput: see the B200 guide and H200 vs B200 vs GB200.
FAQ
Is the H200 faster than the H100?
Only where memory matters. Tensor-core throughput is identical in NVIDIA's tables. NVIDIA reports up to 28% better Llama 2 70B performance at the same 700 W in MLPerf Inference v4.0, and up to 1.9x on its product page with larger batches.
How much more memory does the H200 have?
141 GB of HBM3e against 80 GB of HBM3 on the H100 SXM, about 1.76x, with 4.8 TB/s against 3.35 TB/s of bandwidth, about 1.43x.
Should I upgrade from H100 to H200?
If your model, KV cache or batch size is limited by 80 GB, yes. If your job is compute-bound or already fits comfortably, the gain is small and the H100 may be cheaper per token.
Do H100 and H200 use the same servers?
NVIDIA says HGX H200 boards are compatible with the hardware and software of HGX H100 systems. Both SXM parts run at up to 700 W.
Is H200 better for training?
Not by much on compute, since the tensor-core numbers match. It can help when memory limits your sharding or micro-batch size.
What about the PCIe versions?
The H100 NVL has 94 GB and the H200 NVL has 141 GB. NVIDIA claims up to 1.7x faster LLM inference for the H200 NVL; see the form factor guide.
Sources
- NVIDIA H100 product page (H100 SXM and NVL specs): https://www.nvidia.com/en-us/data-center/h100/
- NVIDIA H200 product page (H200 SXM and NVL specs, vendor claims with conditions): https://www.nvidia.com/en-us/data-center/h200/
- NVIDIA press release, NVIDIA Supercharges Hopper (HGX H200 compatibility with HGX H100): https://nvidianews.nvidia.com/news/nvidia-supercharges-hopper-the-worlds-leading-ai-computing-platform
- NVIDIA technical blog, H200 and TensorRT-LLM set MLPerf LLM inference records (March 27, 2024): https://developer.nvidia.com/blog/nvidia-h200-tensor-core-gpus-and-nvidia-tensorrt-llm-set-mlperf-llm-inference-records
- NVIDIA technical blog, Deploying H200 NVL at scale (H200 NVL vs H100 NVL claims): https://developer.nvidia.com/blog/deploying-nvidia-h200-nvl-at-scale-with-new-enterprise-reference-architecture