H100 and H200 each ship in two main forms: SXM modules that sit on an HGX board and talk over NVSwitch at 900 GB/s, and dual-slot PCIe cards (NVL for the current generation) that are air-cooled, draw 350 to 600 W and link in pairs or fours through NVLink bridges. The silicon is the same Hopper architecture; what changes is memory, power limit and how GPUs connect to each other.
This guide covers:
- What SXM, NVL and PCIe mean for the H100 and the H200
- Side-by-side specs from NVIDIA's own pages
- How NVLink works in each form factor, and why it matters for tensor parallelism
- Which form factor suits inference, training and enterprise racks
- How to read the rental options for each
TL;DR
- SXM is the full-power module: up to 700 W, NVLink at 900 GB/s per GPU through NVSwitch, in HGX boards of 4 or 8 GPUs. It is the form factor for training and for large tensor-parallel inference.
- NVL is the air-cooled PCIe card with an NVLink bridge. The H100 NVL has 94 GB of HBM3 and a 600 GB/s bridge; the H200 NVL has 141 GB of HBM3e and a 900 GB/s two-way bridge or a four-way NVLink option.
- The original H100 PCIe is the lowest-spec variant: 80 GB of HBM2e at about 2 TB/s and 350 W, according to NVIDIA's Hopper architecture write-up. NVIDIA's current H100 product page lists the NVL as its PCIe-based model.
- Verdict: choose SXM when GPUs need to talk to each other constantly. Choose NVL when you need an air-cooled, standard server and your model fits on one to four GPUs.
For the broader landscape, see our datacenter GPU overview.
What do SXM, PCIe and NVL mean?
SXM is NVIDIA's mezzanine module format. The GPU is not a card you slot into a PCIe connector; it mounts on a baseboard (HGX) with its own power delivery and connects to the other GPUs on the board through NVSwitch chips. NVIDIA's H200 page lists the server options for the SXM version as HGX H200 partner and NVIDIA-Certified Systems with 4 or 8 GPUs.
PCIe is the familiar add-in card format. It fits a standard server slot and is air-cooled, but its power budget is lower than an SXM module's, and GPU to GPU traffic goes over PCIe unless an NVLink bridge is fitted.
NVL is NVIDIA's name for the PCIe-form-factor Hopper cards that ship with NVLink bridges: the H100 NVL and the H200 NVL. NVIDIA's H100 page lists the H100 NVL in a dual-slot, air-cooled PCIe format. The H200 page lists the NVL with a 2-way or 4-way bridge and "MGX H200 NVL partner and NVIDIA-Certified Systems, up to 8 GPUs".
One naming caution. Do not confuse an H100 NVL or H200 NVL (PCIe cards with bridges) with the NVL72 rack-scale systems in the Blackwell generation. That is a different product with a similar suffix; see HGX vs DGX vs NVL72 for the difference.
Spec comparison
Figures are from NVIDIA's H100 and H200 product pages. NVIDIA quotes tensor-core peaks with sparsity; the table shows those as published and a dense figure calculated by halving them (our arithmetic).
| Spec | H100 SXM | H100 NVL | H200 SXM | H200 NVL |
|---|---|---|---|---|
| GPU memory | 80 GB HBM3 | 94 GB HBM3 | 141 GB HBM3e | 141 GB HBM3e |
| Memory bandwidth | 3.35 TB/s | 3.9 TB/s | 4.8 TB/s | 4.8 TB/s |
| FP8 tensor (sparsity, as published) | 3,958 TFLOPS | 3,341 TFLOPS | 3,958 TFLOPS | 3,341 TFLOPS |
| FP8 tensor (dense, halved) | About 1,979 | About 1,670 | About 1,979 | About 1,670 |
| BF16 tensor (sparsity, as published) | 1,979 TFLOPS | 1,671 TFLOPS | 1,979 TFLOPS | 1,671 TFLOPS |
| Max power | Up to 700 W | 350 to 400 W | Up to 700 W | Up to 600 W |
| NVLink | 900 GB/s | 600 GB/s (bridge) | 900 GB/s | 900 GB/s per GPU (2- or 4-way bridge) |
| PCIe | Gen 5, 128 GB/s | Gen 5, 128 GB/s | Gen 5, 128 GB/s | Gen 5, 128 GB/s |
| MIG | Up to 7 at 10 GB | Up to 7 at 12 GB | Up to 7 at 18 GB | Up to 7 at 16.5 GB |
| Form factor | SXM | PCIe, dual-slot, air-cooled | SXM | PCIe, dual-slot, air-cooled |
| NVIDIA AI Enterprise | Add-on | Included, 5 years | Add-on | Included, 5 years |
Three things stand out. The NVL variants have lower compute than SXM (about 84% of the SXM peak on the tensor-core rows, from NVIDIA's numbers) because of the lower power limit. The H100 NVL has more memory and bandwidth than the H100 SXM (94 GB and 3.9 TB/s against 80 GB and 3.35 TB/s). And the H200 NVL matches the H200 SXM on memory and bandwidth despite its lower power cap, which makes it a strong inference card.
NVIDIA marks the H200 spec figures as preliminary and subject to change on its page, so confirm against the datasheet for the SKU in your quote.
What about the older H100 PCIe?
NVIDIA's Hopper architecture write-up describes the original H100 PCIe differently from the NVL: 114 SMs against 132 on the SXM5, 80 GB of HBM2e at about 2 TB/s against HBM3 on the SXM5, 350 W against 700 W, and a peak FP8 tensor rate of 1,600 TFLOPS dense (3,200 sparse) against 2,000 dense (4,000 sparse). That card is the weakest of the group, and the H100 NVL superseded it with HBM3 and more capacity. When a listing says "H100 PCIe", check whether it means this 80 GB HBM2e card or the 94 GB NVL.
How NVLink differs across form factors
NVLink is the GPU to GPU link, and it is the reason the form factors are not interchangeable.
- SXM. Each H100 has 18 fourth-generation NVLink links, 900 GB/s total, per NVIDIA's Hopper write-up, which also says that is 7x the bandwidth of PCIe Gen 5. On an HGX board NVSwitch chips connect every GPU to every other GPU, so an 8-GPU node behaves like one pool with uniform bandwidth. Our glossary covers the difference in NVLink vs PCIe, and the what is NVLink guide goes through NVSwitch in detail.
- H100 NVL. A bridge links two cards at 600 GB/s. Beyond the pair, traffic uses PCIe Gen 5 at 128 GB/s.
- H200 NVL. NVIDIA's technical blog describes a two-way bridge at 900 GB/s of GPU to GPU bandwidth (50% more than the H100 NVL, and 7x PCIe Gen 5) and a four-way NVLink option at up to 1.8 TB/s that pools 564 GB of HBM3e across four GPUs, 3x the H100 NVL's two-way pool.
Why does this matter? Tensor parallelism splits each layer across GPUs and needs an all-reduce on every layer. That traffic is bandwidth and latency sensitive, so a pair or four-way bridge works for a model that needs two to four GPUs, but an 8-way tensor-parallel job over PCIe between bridged groups will be limited by the slower hop. For an 8-way job, NVSwitch on an SXM board is the clean answer.
Performance: what NVIDIA publishes
All figures are NVIDIA's claims with the conditions it states.
- H200 SXM vs H100 SXM. NVIDIA's H200 page claims 1.9x faster Llama2 70B inference (throughput, input length 2K, output length 128, batch size 32 on the H200 against 8 on the H100 per GPU), and 1.6x faster GPT-3 175B inference (8x H200 against 8x H100, input length 80, output 200, batch size 128 against 64). Note that the batch sizes differ between the two sides of each comparison, so these are not like-for-like per-batch figures.
- H200 NVL vs H100 NVL. NVIDIA's technical blog claims up to 1.7x faster LLM inference and 1.3x more HPC performance, without listing the workload or test conditions.
- H100 NVL vs A100. NVIDIA's H100 page says servers with H100 NVL "increase Llama 2 70B performance up to 5x over NVIDIA A100 systems", with no test conditions listed on the page.
- H100 vs A100, training. The H100 page claims up to 4x higher GPT-3 training performance, labelled as projected and subject to change, on a cluster comparison that uses a different network on each side.
None of these compare SXM against NVL on the same workload, and we have not found an NVIDIA table that does. If the choice between form factors matters to your budget, benchmark your own model on both.
Infrastructure needs
| SXM (HGX) | NVL / PCIe | |
|---|---|---|
| Power per GPU | Up to 700 W | 350 to 600 W |
| Cooling | Set by the OEM system design; ask the vendor | Air-cooled, dual-slot |
| Chassis | HGX baseboard with 4 or 8 GPUs | MGX or standard PCIe server, up to 8 GPUs |
| Host CPU to GPU | PCIe Gen 5 | PCIe Gen 5 |
| Network guidance | Depends on the OEM design | NVIDIA's H200 NVL reference architecture recommends one BlueField-3 SuperNIC at 400 Gb/s for every two GPUs |
NVIDIA's H200 NVL enterprise reference architecture uses a "PCIe Optimized 2-8-5" configuration: 2 CPUs, 8 GPUs and 5 network adapters. That is designed to drop into existing air-cooled enterprise racks without the power and cooling upgrade an HGX system needs. The NVL cards also include a five-year NVIDIA AI Enterprise subscription per NVIDIA, whereas it is an add-on for SXM.
How to choose
Choose SXM (H100 or H200) when:
- You train models, or run multi-node jobs. NVSwitch bandwidth and the 700 W power limit show up in step time.
- You serve a large model with tensor parallelism across 4 or 8 GPUs.
- You want the highest compute per GPU in the Hopper family.
Choose H100 NVL or H200 NVL when:
- You serve models that fit on one to four GPUs and you value memory capacity over peak compute. The H200 NVL's 141 GB at 4.8 TB/s is the notable spec.
- You need to deploy in a standard air-cooled rack.
- Your serving is memory-bound decode rather than compute-bound prefill.
Choose H200 over H100 when memory is your limit. The jump from 80 GB to 141 GB, and from 3.35 to 4.8 TB/s on SXM, is the whole difference between the two chips; see the H100 vs H200 comparison. If you are choosing between Hopper and Blackwell, our H200 vs B200 vs GB200 guide covers it.
Avoid assuming a "PCIe" listing is the same as the NVL. Ask for the memory type and size.
Cost: how to compare
We do not type rental prices because they move; the box below shows the live from-price per GPU-hour for each chip. To compare form factors, use tokens per GPU-hour: measured tokens per second times 3,600. Then divide the live hourly price by that number for cost per token. NVL cards can win on cost when the model fits and the workload is memory-bound, and SXM wins when scaling efficiency across 8 GPUs is the constraint. Measure both on your workload.
Rent today
Aquanode manages and optimizes GPUs for training and inference workloads, and you can rent the GPUs in the box below on demand. The H100 NVL, H100 and H200 listings appear with their current from-price, or "None right now" if no offer is available.
Details on each: H100 NVL, H100, H200, and /pricing. To see the whole field, use the GPU index.
What's next
Blackwell replaces both the SXM and NVL Hopper parts at the top end; start with the B200 guide or the Rubin vs Blackwell vs Hopper overview. For the Hopper variant with an Arm CPU on the module, see the GH200 guide.
FAQ
What is the difference between H100 SXM and PCIe?
SXM is a module on an HGX board with up to 700 W and 900 GB/s NVLink. The PCIe-format cards run at lower power (350 to 400 W for the H100 NVL), connect with a 600 GB/s bridge rather than NVSwitch, and are air-cooled.
What is the H100 NVL?
NVIDIA's PCIe-format H100 with 94 GB of HBM3, 3.9 TB/s of bandwidth, a 350 to 400 W configurable power limit and a 600 GB/s NVLink bridge between two cards. NVIDIA's page cites 188 GB for a pair, which is two times 94 GB.
What is the H200 NVL?
The PCIe-format H200: 141 GB of HBM3e at 4.8 TB/s, up to 600 W, in a dual-slot air-cooled card. It supports a 2-way bridge at 900 GB/s and a 4-way NVLink option, and includes a five-year NVIDIA AI Enterprise subscription.
Is the H200 NVL slower than the H200 SXM?
On tensor-core peaks, yes: NVIDIA lists 3,341 TFLOPS FP8 with sparsity for the NVL against 3,958 for the SXM. Memory capacity and bandwidth are identical at 141 GB and 4.8 TB/s.
Can I run 8-way tensor parallelism on NVL cards?
You can, but only groups bridged together get NVLink bandwidth, so the slowest hop between groups sets the pace. NVSwitch on an SXM board gives uniform bandwidth across all 8 GPUs.
Which is best for inference?
For models that fit on one to four GPUs, the H200 NVL and H200 SXM have the best memory specs in the Hopper family. For very large models split across 8 GPUs, use SXM.
Sources
- NVIDIA H100 product page (SXM and NVL spec table, claims): https://www.nvidia.com/en-us/data-center/h100/
- NVIDIA H200 product page (SXM and NVL spec table, claims): https://www.nvidia.com/en-us/data-center/h200/
- NVIDIA technical blog, Hopper architecture in depth (SXM5 vs PCIe table, NVLink): https://developer.nvidia.com/blog/nvidia-hopper-architecture-in-depth/
- NVIDIA technical blog, Deploying H200 NVL at scale with a new enterprise reference architecture: https://developer.nvidia.com/blog/deploying-nvidia-h200-nvl-at-scale-with-new-enterprise-reference-architecture