The NVIDIA H20 is a Hopper-generation datacenter GPU that NVIDIA built for the China market to fit US export rules: it keeps a large 96 GB of HBM3 memory with about 4 TB/s of bandwidth, but its compute is cut to a small fraction of an H100. The US government required a license for H20 exports to China in April 2025, and as of October 2026 it is not a chip you can buy or rent through normal channels.
This guide covers:
- What the H20 is and why it exists
- Its reported specs, and why NVIDIA's own spec sheet is not available to compare against
- A dated timeline of the export-license history
- What the H20 is actually good at (and bad at) compared with an H100 or H200
- What to run if you wanted an H20-class workload today
TL;DR
- H20 is a cut-down Hopper: reported 96 GB HBM3, about 4 TB/s of memory bandwidth, and roughly 148 TFLOPS FP16 and 296 TFLOPS FP8 (dense), according to Tom's Hardware's spec table. NVIDIA has not published a public datasheet that we could find.
- By our arithmetic from those figures, the H20 offers about 15% of an H100 SXM's dense FP8 compute, while keeping more memory and more memory bandwidth than the H100.
- Export timeline: license required in April 2025, a $4.5 billion H20 charge that quarter, limited licenses in August 2025 producing about $60 million of H20 revenue, and Beijing discouraging purchases in the same period. The US shifted to licensing the H200 for China in December 2025.
- Verdict: the H20 was designed around a policy limit, not a workload. For LLM serving that is memory-bound, the memory spec is attractive; for training or compute-heavy inference, an H100 or H200 is the practical choice, and both are available to rent below.
For where H20 sits among all datacenter accelerators, see our datacenter GPU overview.
What is the NVIDIA H20?
H20 is one of the China-specific parts NVIDIA designed after the US tightened export controls on advanced AI chips; SemiAnalysis wrote about the family in November 2023. NVIDIA's approach, as the reported specs suggest, was to build a Hopper-architecture chip with reduced compute while keeping a large memory system. NVIDIA's own filing later describes the April 2025 license requirement as covering the H20 and any circuit matching its memory or interconnect bandwidth, so memory and interconnect are part of the rule too.
That is why the H20 looks unusual on paper. The Hopper architecture is the same one inside the H100 and H200 (see our GH200 guide for another Hopper variant), but fewer of its tensor-core resources are enabled, so peak throughput drops while the HBM stack stays large.
H20 specs (reported)
NVIDIA does not publish a public H20 datasheet that we could find, so the H20 figures below come from Tom's Hardware's HGX H20 spec table and should be treated as reported, not vendor-confirmed. The H100 and H200 columns come from NVIDIA's own product pages, where tensor-core peaks are quoted with sparsity; we halve them to get the dense figures, which is the basis Tom's Hardware uses for the H20.
| Spec | H20 (reported) | H100 SXM | H200 SXM |
|---|---|---|---|
| Architecture | Hopper | Hopper | Hopper |
| GPU memory | 96 GB HBM3 | 80 GB HBM3 | 141 GB HBM3e |
| Memory bandwidth | About 4.0 TB/s | 3.35 TB/s | 4.8 TB/s |
| FP16 / BF16 tensor, dense | About 148 TFLOPS | About 990 TFLOPS | About 990 TFLOPS |
| FP8 tensor, dense | About 296 TFLOPS | About 1,979 TFLOPS | About 1,979 TFLOPS |
| Max power (TDP) | 400 W (reported) | Up to 700 W | Up to 700 W |
| NVLink | 900 GB/s (reported) | 900 GB/s | 900 GB/s |
How to read the table:
- Memory is the headline. At 96 GB and about 4 TB/s, the H20 has more capacity than an H100 and more bandwidth than the H100's 3.35 TB/s. It trails the H200 on both.
- Compute is where it was cut. Dividing the reported 296 TFLOPS FP8 by the H100 SXM's 1,979 TFLOPS dense FP8 (NVIDIA's 3,958 with sparsity, halved) gives about 15%. That ratio is our calculation from two sources, not a vendor figure.
- NVLink. Tom's Hardware lists 900 GB/s of NVLink on an 8-way HGX board. We have not confirmed that against an NVIDIA document.
- Variants. Some databases list a 141 GB HBM3e H20. We found no NVIDIA or major-press source for it and treat the 96 GB HBM3 part as the H20.
If you need H20 numbers for a procurement decision, ask the system vendor for the datasheet that ships with the exact SKU.
Why a compute-cut chip still has a use
LLM inference splits into two phases. Prefill processes the prompt and is compute-bound. Decode generates tokens one at a time and is dominated by reading weights and the KV cache from memory, so it is memory-bandwidth-bound. A chip with weak compute and strong memory is less bad at decode than its peak TFLOPS suggests.
That is the argument for H20 in serving: capacity for large models and long contexts (see the KV cache entry in our glossary), and enough bandwidth to keep decode moving. SemiAnalysis reported in November 2023, around the time the chips were first described, that one of the China-specific GPUs could be over 20% faster than the H100 in LLM inference; the details were behind the publication's paywall, so we cite that as a reported claim only, and it applies to specific inference cases, not training.
Where it falls short:
- Training. Training is compute-bound at scale, and 15% of an H100's FP8 peak means many more GPUs, more power and more network for the same job.
- Prefill-heavy serving. Long prompts and large batches hit the compute ceiling first.
- Anything needing FP8 throughput. The Transformer Engine gains come from FP8 tensor throughput, which is exactly what is reduced.
Export status timeline
Dates below come from NVIDIA's SEC filings unless marked as press reporting.
- April 2025. The US government told NVIDIA it requires a license to export the H20, and any circuit matching its memory or interconnect bandwidth, to China (including Hong Kong and Macau) and D:5 countries. (NVIDIA Form 10-K for fiscal 2026.)
- First quarter of fiscal 2026. NVIDIA took a $4.5 billion charge associated with the H20 for excess inventory and purchase obligations. (10-K.)
- August 2025. The US government granted licenses allowing certain H20 products to ship to certain China-based customers. NVIDIA reports generating approximately $60 million of H20 revenue under those licenses, and says officials expressed an expectation of receiving 15% or more of the revenue from licensed sales, without publishing a regulation to that effect. (10-K and 10-Q.)
- August 2025. Press reporting: Reuters reported that Chinese authorities summoned companies including Tencent and ByteDance over H20 purchases and asked them to explain why they bought NVIDIA chips when domestic suppliers were available, and that officials had information-security concerns. Reuters' sources said the companies had not been ordered to stop buying. NVIDIA said the H20 was "not a military product or for government infrastructure." (Reuters via Al Jazeera, August 12, 2025.)
- December 8, 2025. Press reporting: President Trump said NVIDIA could ship H200 products to approved customers in China with 25% going to the US government, and that the Commerce Department was finalizing details. (The Register, December 9, 2025.)
- February 2026. NVIDIA's filing says the US government granted a license allowing small amounts of H200 products to ship to specific China-based customers, with the units passing an inspection process in the US, and that it had generated no revenue under the H200 program as of that filing. (10-K.)
- March 2026. CNBC reported that Jensen Huang told reporters at GTC, "We have received purchase orders, and we're in the process of restarting our manufacturing," as NVIDIA prepared to sell H200 processors to some customers in China. (CNBC, March 17, 2026.)
- July 2026 filing. NVIDIA's latest quarterly report (10-Q, quarter ended July 26, 2026) says H200 licenses exist but sales were restricted by the PRC government and NVIDIA has been unable to sell all the products it is licensed for, that it is "effectively foreclosed" from competing in China's data center market, and that it needs a data center system that meets the approval of both the US and Chinese governments.
What we did not find: any 2026 source stating that the H20 has resumed volume shipments to China. Our reading is that US policy attention moved to the H200, and that Chinese buyers were already being steered toward domestic chips. Treat the H20's current status as "licensed in principle, commercially dormant" only as a summary of the sources above, and check the latest NVIDIA filing before relying on it.
How the H20 compares with domestic alternatives
Chinese buyers have the alternative of domestic accelerators. For a neutral comparison of the leading domestic line against NVIDIA, read our Huawei Ascend vs NVIDIA guide. We do not repeat those numbers here because the comparison needs its own sourcing.
When to choose what
| Situation | Better fit |
|---|---|
| Training any model beyond fine-tuning scale | H100 or H200, or Blackwell |
| Decode-heavy serving of a large model, memory-bound | H200 (141 GB, 4.8 TB/s) |
| Long-context serving where KV cache dominates | H200, or GH200 with its extra CPU memory |
| Prefill-heavy or batch inference | H100 or H200 |
| Need the newest FP4 and FP8 throughput | Blackwell: see the B200 guide |
If your reason for looking at the H20 is its memory-to-compute ratio, the H200 is the closer match in practice: more memory, more bandwidth, and about 6.7 times the dense FP8 compute by the same arithmetic as above (1,979 divided by 296).
Cost: how to compare
We do not quote a rental price here; it changes by the hour and the box below shows the live figure. To compare chips fairly, use tokens per GPU-hour from a measured or published throughput: tokens per second times 3,600. Then multiply by the live hourly price below to get cost per million tokens.
Be careful with spec-sheet shortcuts. A chip with 15% of the compute does not need to be priced at 15% of an H100 to be good value, because decode is memory-bound; but you cannot assume it either. The only reliable number is throughput on your own model and context length.
Rent today
The H20 is not offered on Aquanode. Aquanode manages and optimizes GPUs for training and inference workloads, and you can rent the GPUs in the box below on demand. The H100 and H200 are the closest available Hopper options.
See the H100 page and H200 page for specs, /pricing for billing, and the H100 vs H200 comparison for a side-by-side.
What's next
Export policy for China moved from the H20 to the H200 in late 2025 and early 2026, and NVIDIA's filings say further China-specific products depend on approval from both governments. For the architecture roadmap beyond Hopper, see Rubin vs Blackwell vs Hopper. We do not date any future China-specific part, because none is confirmed in the sources we reviewed.
FAQ
Is the NVIDIA H20 banned?
It is not banned outright. Since April 2025 the US requires a license to export it to China, and licenses were granted for certain customers in August 2025. NVIDIA reported about $60 million of H20 revenue under those licenses. Chinese authorities also discouraged purchases.
What are the NVIDIA H20 specs?
Reported specs are 96 GB HBM3, about 4 TB/s memory bandwidth, about 148 TFLOPS FP16 and about 296 TFLOPS FP8 (dense), and a 400 W TDP. NVIDIA has no public datasheet we could find, so these come from Tom's Hardware's spec table.
Is the H20 faster than the H100?
Not in compute: by the reported figures it has about 15% of an H100's dense FP8 throughput. For memory-bound LLM decode it can look better than its TFLOPS suggest, because it has more memory than an H100 and more bandwidth than the H100's 3.35 TB/s.
Can I rent an H20?
Not on Aquanode. The box above lists what is available. For a similar Hopper architecture with far more compute, rent an H100 or H200.
Why did NVIDIA take a $4.5 billion charge on the H20?
NVIDIA's filing says the charge, taken in the first quarter of fiscal 2026, covered excess H20 inventory and purchase obligations as demand diminished after the April 2025 license requirement.
What replaced the H20 for China?
US policy moved toward licensing the H200 in December 2025, and a February 2026 filing describes a license for small amounts of H200 products. NVIDIA reported no H200 revenue under that program as of its filing.
Sources
- NVIDIA Form 10-K for fiscal 2026 (H20 license, $4.5 billion charge, August 2025 licenses, February 2026 H200 license): https://www.sec.gov/Archives/edgar/data/1045810/000104581026000021/nvda-20260125.htm
- NVIDIA Form 10-Q, quarter ended July 26, 2026 (export control language): https://www.sec.gov/Archives/edgar/data/0001045810/000104581026000075/nvda-20260726.htm
- Al Jazeera, citing Reuters, China raises concerns over Nvidia's H20 chips (August 12, 2025): https://www.aljazeera.com/economy/2025/8/12/china-raises-concerns-over-nvidias-h20-chips-with-local-firms-report
- The Register, Trump says Nvidia can sell H200s to China (December 9, 2025): https://www.theregister.com/2025/12/09/trump_gpu_export_ban_reversal/
- CNBC, Jensen Huang says Nvidia has received orders from China (March 17, 2026): https://www.cnbc.com/2026/03/17/nvidia-ceo-jensen-huang-says-chipmaker-has-received-orders-from-china.html
- TechCrunch, US government imposes license requirement on Nvidia H20 exports (April 15, 2025): https://techcrunch.com/2025/04/15/us-government-imposes-license-requirement-on-nvidia-h20-exports
- SemiAnalysis, Nvidia's new China AI chips circumvent US restrictions (November 9, 2023; paywalled): https://newsletter.semianalysis.com/p/nvidias-new-china-ai-chips-circumvent
- Tom's Hardware, The tale of Nvidia's HGX H20 (96 GB HBM3, 4.0 TB/s, 148 TFLOPS FP16, 296 TFLOPS FP8, 400 W, 900 GB/s NVLink; reported specs): https://www.tomshardware.com/pc-components/gpus/the-tale-of-nvidias-hgx-h20-how-an-ai-gpu-became-a-political-lightning-rod
- NVIDIA H100 product page (H100 SXM figures, quoted with sparsity): https://www.nvidia.com/en-us/data-center/h100/
- NVIDIA H200 product page (H200 SXM figures): https://www.nvidia.com/en-us/data-center/h200/