NVIDIA H200: Specs, VRAM and Memory Guide (2026)

Back
Team Aquanode

Team Aquanode

Sarthak Vaish

Updated OCTOBER 8, 2026Published OCTOBER 8, 2026

The NVIDIA H200 is a Hopper GPU with 141 GB of HBM3e memory and 4.8 TB/s of memory bandwidth, which is 76% more capacity and 43% more bandwidth than the H100 SXM's 80 GB and 3.35 TB/s. Its compute is the same as the H100's, so the H200 wins where memory is the limit (large models, long contexts, big batches) and ties where compute is.

This guide covers:

  • H200 specs, including VRAM, memory bandwidth, NVLink and power
  • What HBM3e changes and how to size a model against 141 GB
  • NVIDIA's published performance claims and MLPerf results, labelled as the vendor's
  • Infrastructure needs and when to pick an H200
  • How to read cost per token

TL;DR

  • Memory is the whole upgrade. H200 SXM: 141 GB HBM3e at 4.8 TB/s. H100 SXM: 80 GB HBM3 at 3.35 TB/s. Tensor-core TFLOPS are identical in NVIDIA's tables.
  • Same power class. The SXM version is configurable up to 700 W, like the H100 SXM, and NVIDIA says HGX H200 boards are compatible with HGX H100 hardware and software.
  • Vendor claims. NVIDIA claims 1.9x faster Llama2 70B inference than the H100 and, in MLPerf Inference v4.0, up to 28% better at the same 700 W (up to 45% in a 1,000 W custom-cooled configuration).
  • Verdict: choose the H200 for inference on 70B-class models, long-context serving and fine-tuning where 80 GB forces you to shard. Stay on the H100 when the model already fits and compute is the limit.

For the full lineup, see our datacenter GPU overview. For a head-to-head, read H100 vs H200.

What is the NVIDIA H200?

NVIDIA announced the H200 in November 2023 as the first GPU with HBM3e. In its press release NVIDIA describes 141 GB of memory at 4.8 terabytes per second, "nearly double the capacity and 2.4x more bandwidth compared with its predecessor, the NVIDIA A100", with availability from cloud and system partners starting in the second quarter of 2024. The same release says HGX H200 boards come in four- and eight-way configurations and are compatible with both the hardware and software of HGX H100 systems.

Architecturally it is the Hopper GPU you already know. The product page lists the same tensor-core throughput as the H100 and the same fourth-generation NVLink. What NVIDIA changed is the memory stack: faster, denser HBM3e in place of HBM3. That is why the H200 is best understood as an H100 that no longer runs out of memory as early.

The H200 ships in two main form factors, SXM and the PCIe-based NVL. This guide covers the chip; the form factors, bridges and rack designs are in H100 and H200: SXM vs NVL vs PCIe.

H200 specs

From NVIDIA's H200 product page (NVIDIA marks these figures as preliminary and subject to change). Tensor-core figures are quoted with sparsity; dense is half, which is our arithmetic.

SpecH200 SXMH200 NVL
GPU memory (VRAM)141 GB HBM3e141 GB HBM3e
Memory bandwidth4.8 TB/s4.8 TB/s
FP8 tensor (with sparsity)3,958 TFLOPS3,341 TFLOPS
BF16 / FP16 tensor (with sparsity)1,979 TFLOPS1,671 TFLOPS
TF32 tensor (with sparsity)989 TFLOPS835 TFLOPS
FP64 / FP64 tensor34 / 67 TFLOPS30 / 60 TFLOPS
Max power (TDP)Up to 700 W (configurable)Up to 600 W (configurable)
NVLink900 GB/s900 GB/s per GPU (2- or 4-way bridge)
PCIeGen 5, 128 GB/sGen 5, 128 GB/s
MIGUp to 7 instances at 18 GBUp to 7 instances at 16.5 GB
Decoders7 NVDEC, 7 JPEG7 NVDEC, 7 JPEG
Confidential computingSupportedSupported
Server optionsHGX H200, 4 or 8 GPUsMGX H200 NVL, up to 8 GPUs

H200 VRAM and memory bandwidth in context

Compared with the H100 SXM (80 GB HBM3, 3.35 TB/s, from NVIDIA's H100 product page), the ratios are our arithmetic: 141 divided by 80 is about 1.76x capacity, and 4.8 divided by 3.35 is about 1.43x bandwidth. NVIDIA's own MLPerf write-up rounds these to roughly 1.8x and 1.4x.

What HBM3e changes

HBM stacks DRAM dies vertically next to the GPU die for far higher bandwidth than conventional memory; see HBM in our glossary. HBM3e is the newer generation, with more capacity and bandwidth per stack. The follow-on generation is covered in HBM3e vs HBM4.

For LLMs the memory system matters in two ways:

  1. Capacity decides what fits. Weights, KV cache and activations all share VRAM.
  2. Bandwidth decides decode speed. Generating each token requires reading the model weights and the KV cache from memory, so single-stream decode is limited by bandwidth, not by TFLOPS.

NVIDIA's MLPerf explanation says the same: the extra memory removes the need for tensor or pipeline parallelism on Llama 2 70B in that benchmark, which cuts communication overhead, and the higher bandwidth relieves memory-bound bottlenecks.

How much model fits in 141 GB?

Weight memory is parameters times bytes per parameter. This is arithmetic, not a measurement, and it ignores the KV cache, activations and framework overhead, which you must budget on top.

Model sizeBF16 weights (2 bytes)FP8 weights (1 byte)Fits in 141 GB?
8B16 GB8 GBYes, with large KV cache room
34B68 GB34 GBYes in either precision
70B140 GB70 GBBF16: no practical room; FP8: yes, with about 70 GB spare
120B240 GB120 GBFP8: tight; use 2 GPUs or lower precision

The 70B row is the important one. In BF16, a 70B model's weights alone take about 140 GB, leaving essentially nothing for the KV cache on a single 141 GB card. In FP8 they take about 70 GB, which fits on an H200 with room for cache, and fits on an 80 GB H100 only barely. That is the practical meaning of "the H200 is the 70B card". For how FP8 works in practice, see Transformer Engine and FP8 and the quantization glossary entry. For KV cache sizing, see KV cache.

Performance: what NVIDIA has published

Every number below is NVIDIA's claim or an MLPerf result. We did not benchmark the H200 ourselves.

Product page claims (vendor's claims):

  • Llama2 70B inference: 1.9x faster than H100 SXM. Conditions: throughput, input length 2K, output length 128, batch size 32 on the H200 against 8 on the H100 per GPU.
  • GPT-3 175B inference: 1.6x faster, 8x H200 SXM against 8x H100 SXM, input length 80, output length 200, batch size 128 against 64.
  • General LLM inference: up to 2x faster than H100 for models such as Llama2.
  • HPC: up to 110x faster time to results than CPUs on the MILC benchmark (HGX H200 4-GPU against a dual Sapphire Rapids 8480).

The batch sizes differ between the two sides in the first two claims. They show what the larger memory allows, not an equal-batch comparison.

MLPerf Inference v4.0 (NVIDIA's blog, results retrieved from MLCommons on March 27, 2024):

  • Llama 2 70B, same 700 W TDP as the H100: up to 28% better inference performance than the H100.
  • Llama 2 70B, H200 in a custom thermal design at a 1,000 W TDP: up to 45% more performance than the H100. The higher power setting added 11% in the server scenario and 14% offline over the 700 W H200, for total speedups of 43% (server) and 45% (offline). The submissions are 4.0-0062 and 4.0-0068.
  • Stable Diffusion XL, 8-GPU HGX H200 at 700 W: 13.8 queries per second (server) and 13.7 samples per second (offline), against 4.9 and 5 for an 8-GPU L40S system.

Notice how much smaller the 28% MLPerf figure is than the 1.9x product page figure. The MLPerf run uses fixed rules and the same configuration on both sides; the product page claim uses larger batches on the H200. Plan around the lower figure unless your workload also gains from a larger batch, in which case you can see more.

NVIDIA's later MLPerf rounds exist but we have not reviewed them for this guide; check MLCommons for the latest.

Infrastructure needs

  • Power. The SXM version is configurable up to 700 W, same as the H100 SXM. The NVL card goes up to 600 W. NVIDIA's MLPerf v4.0 1,000 W result used a custom thermal design, so it is not a standard configuration.
  • Drop-in compatibility. NVIDIA says HGX H200 boards are compatible with HGX H100 hardware and software, which makes upgrading an existing 8-GPU H100 server design straightforward in principle. Ask your system vendor to confirm for a specific chassis.
  • Interconnect. 900 GB/s NVLink on SXM; see what is NVLink for how NVSwitch ties 8 GPUs together. The NVL card supports a 2-way or 4-way bridge.
  • Software. NVIDIA says HGX H200 is compatible with HGX H100 software, so existing H100 stacks should carry over. MIG partitions are larger (up to 7 instances at 18 GB on SXM).
  • Confidential computing is supported, per NVIDIA's table.

When to choose the H200

Choose the H200 when:

  • You serve 70B-class models. FP8 weights fit with room for the KV cache on one GPU, avoiding tensor parallelism.
  • You serve long contexts or large batches. KV cache grows with both, and 141 GB gives more room before you hit the memory wall.
  • Decode is your bottleneck. The 43% bandwidth gain translates directly into faster token generation for memory-bound decode, in principle.
  • You fine-tune larger models and want to keep a replica on one GPU.

Choose the H100 when:

  • Your model fits comfortably in 80 GB and the job is compute-bound, such as prefill-heavy serving or training with small models, where the two chips have the same tensor-core throughput.
  • The price difference per GPU-hour is large relative to your measured gain.

Choose something newer when you need FP4 or much more bandwidth: the B200 guide covers Blackwell, and H200 vs B200 vs GB200 compares the generations. For the CPU-attached variant with additional memory, see the GH200 guide.

Cost: how to compare

We do not print an hourly price because it changes; the live box below shows the current from-price. To compare the H200 with the H100:

  1. Measure tokens per second on your model on each GPU.
  2. Multiply by 3,600 for tokens per GPU-hour.
  3. Divide the live hourly price by that number, then by one million, to get cost per million tokens.

Because the H200's gain is largest when the model needs the extra memory, run the test at your real batch size and context length. A test at a small batch can show no gain at all.

Rent today

Aquanode manages and optimizes GPUs for training and inference workloads, and you can rent the GPUs in the box below on demand.

Specs and availability are on the H200 page and H100 page; the live side-by-side shows both together. Billing details are at /pricing.

What's next

The Blackwell generation raises memory and bandwidth again; see the B200 guide and the Rubin vs Blackwell vs Hopper overview. Browse every chip on the GPU index.

FAQ

How much VRAM does the H200 have?

141 GB of HBM3e on both the SXM and NVL versions, per NVIDIA's product page.

What is the H200's memory bandwidth?

4.8 TB/s, against 3.35 TB/s for the H100 SXM.

Is the H200 faster than the H100?

For compute-bound work, no: NVIDIA lists identical tensor-core throughput. For memory-bound LLM inference, NVIDIA reports up to 28% better Llama 2 70B performance in MLPerf Inference v4.0 at the same 700 W, and claims up to 1.9x with larger batches on its product page.

Can an H200 run a 70B model on one GPU?

In FP8, weights take about 70 GB (our arithmetic), which fits in 141 GB with room for KV cache. In BF16 the weights alone take about 140 GB, which leaves no practical room.

How much power does the H200 use?

Up to 700 W (configurable) for SXM and up to 600 W (configurable) for the NVL card, per NVIDIA.

What is the difference between H200 SXM and NVL?

Both have 141 GB and 4.8 TB/s. SXM has higher tensor-core peaks and a higher power limit; NVL is a dual-slot air-cooled PCIe card. See the form factor guide.

Sources

#datacenter gpu#nvidia hopper#h200#hbm3e#h200 vram#llm inference

Submit the job. Everything after that is ours.

Sign up in 60 seconds. Pay for the GPU minutes you actually use.

© 2026 Aquanode. All rights reserved.

All trademarks, logos and brand names are the property of their respective owners.