Qualcomm AI200 is a rack-scale inference accelerator announced on October 27, 2025, with 768 GB of LPDDR memory per card and direct liquid cooling. Qualcomm said it would be available in 2026 and the AI250 in 2027, but we found no public confirmation that AI200 racks have shipped, and Qualcomm has not published compute performance, so it cannot yet be compared with NVIDIA on speed.
TL;DR
- Announced: October 27, 2025. Two products, AI200 (2026) and AI250 (2027), sold as full liquid-cooled racks.
- AI200: 768 GB of LPDDR per card, PCIe for scale-up, Ethernet for scale-out. The launch coverage says the racks aim for 160 kW each (ServeTheHome). We could not open Qualcomm's own product pages, so we give no later rack figures.
- AI250: a near-memory computing design with "greater than 10x higher effective memory bandwidth", against an unstated baseline (Qualcomm's claim).
- Status: Humain in Saudi Arabia is the named customer, targeting 200 MW of AI200 and AI250 racks starting in 2026. We found no report that deliveries have started.
- Not published: compute (TOPS or FLOPS), cards per rack, absolute memory bandwidth of the AI200, price, and any benchmark or MLPerf result.
- Verdict: a capacity-first inference design to watch. Today there is no data to say it beats NVIDIA; use NVIDIA GPUs for inference you need to run now.
What Qualcomm announced
Per coverage of the October 27, 2025 announcement, Qualcomm launched the AI200 and AI250 as rack-scale systems for large language and multimodal inference. Both use direct liquid cooling, PCIe for scale-up, Ethernet for scale-out and confidential computing. The AI200 uses LPDDR memory, which gives high capacity at lower cost than HBM but much lower bandwidth, a trade-off that matters for token generation (see our HBM glossary entry).
ServeTheHome reports the AI200 at 768 GB of LPDDR per card, with racks aiming for 160 kW. It reports no AI250 memory capacity and no rack-level memory total. The AI250 claim is more than 10x higher effective memory bandwidth, which is Qualcomm's claim as relayed by that outlet.
Qualcomm's own product pages are scripted and did not load for us, so we use no figures from them. A June 2026 Qualcomm release unveiled a wider "Dragonfly" data center roadmap; we could not open its full text, so we do not repeat details.
Spec table
| Spec | Qualcomm AI200 | Qualcomm AI250 | NVIDIA H200 | NVIDIA DGX B200 (8 GPUs) |
|---|---|---|---|---|
| Memory per card | 768 GB LPDDR | not published | 141 GB HBM3e | 180 GB average (1,440 GB total / 8) |
| Memory bandwidth | not published | not published | 4.8 TB/s | 64 TB/s total |
| Compute | not published | not published | not covered here | 72 PFLOPS FP8, 144 PFLOPS FP4 (sparse) |
| Rack power | about 160 kW (launch target) | not published | not applicable | about 14.3 kW per 8-GPU system |
| Cooling | direct liquid | direct liquid | air or liquid by system | per NVIDIA |
| Timing | 2026 | 2027 | shipping | shipping |
The NVIDIA B200 memory per GPU is our arithmetic (total divided by eight), not a figure printed on NVIDIA's page. The DGX B200 is an 8-GPU system, not a rack, so its power is not comparable with a full Qualcomm rack. For a rack-scale NVIDIA reference see the GB200 NVL72 guide.
Why LPDDR changes the picture
Inference has two phases. Prefill is compute-heavy. Decode, which generates tokens one by one, is mostly limited by how fast weights and the KV cache can be read from memory. HBM gives an H200 4.8 TB/s per GPU. LPDDR gives far more capacity per dollar and watt but lower bandwidth, so Qualcomm's AI200 pitch is about holding very large models and long contexts, and the AI250 is the answer to the bandwidth problem. Whether the AI250's 10x claim closes the gap depends on a baseline Qualcomm has not stated. Our KV cache glossary entry explains why long contexts need so much memory.
What has shipped
We want to be plain about this.
- Announced: October 27, 2025.
- Customer: Humain (Saudi Arabia) said it is targeting 200 MW of AI200 and AI250 racks starting in 2026, per Qualcomm. The announcement said deployments would start in 2026.
- Shipping: we found no 2026 report confirming that AI200 racks have been delivered or are in production use. Qualcomm published a post about its AI200 rack and infrastructure management software in March 2026, but we could not read it.
- September 8, 2026: Qualcomm announced a multi-generation collaboration with Amazon on customized inference silicon. That release does not name the AI200 or AI250.
- Independent benchmarks: none found.
If you need current status, check Qualcomm's investor materials and newsroom; this field changes quickly.
Software and deployment
Qualcomm describes the racks as arriving with its own rack and infrastructure management software, and confidential computing support for running models on shared hardware. We did not find a published model support list, framework compatibility statement or serving benchmarks. That matters because NVIDIA's advantage in inference is as much software (CUDA, TensorRT-LLM, vLLM and SGLang support from day one) as silicon. Ask any vendor in this category which models, precisions and serving frameworks are validated, and what a migration from a CUDA-based stack would involve. Our quantization glossary entry explains why precision support decides whether a published capacity number is usable.
Infrastructure needs
Plan for a liquid-cooled rack with facility power in the 160 kW class, Ethernet fabric and PCIe-connected cards. This is a datacenter deployment, not a card you drop in a server. By comparison, our HGX vs DGX vs NVL72 guide explains the NVIDIA options at server and rack level.
When to choose which
- Need inference capacity this quarter: NVIDIA. Qualcomm has no public availability beyond a named anchor customer.
- Planning large-scale, memory-heavy inference for 2027 and later, with a facility ready for liquid cooling: ask Qualcomm for a benchmark on your model, with batch size and precision stated, and compare against H200 and B200.
- Looking at other alternatives: Cerebras, Etched, Tenstorrent and TPU. The full map is in the datacenter GPU overview.
Cost
Qualcomm has not published pricing, and it has said its claims are about total cost of ownership without a stated baseline in the pages we read, so we do not convert anything to cost per token. For NVIDIA GPUs, measure tokens per second on your model, multiply by 3,600, and compare with the live hourly price below.
Rent today
Qualcomm AI200 is not something you can rent on demand today. Aquanode manages and optimizes GPUs for training and inference workloads, and you can rent the NVIDIA GPUs below on demand to benchmark your own model.
What's next
Qualcomm's own timing is AI250 in 2027, with a roadmap that its June 2026 release describes as annual. Wait for a reproducible benchmark and a delivery confirmation. On the NVIDIA side, see the Rubin guide.
FAQ
When was the Qualcomm AI200 announced?
October 27, 2025, together with the AI250.
How much memory does the AI200 have?
768 GB of LPDDR per card, per Qualcomm. We found no rack-level memory total from a source we could open.
Has the AI200 shipped?
We found no confirmation. Qualcomm said deployments with Humain would start in 2026.
How does it compare with an H100 or H200?
We cannot say on speed: Qualcomm has not published compute or bandwidth figures for the AI200. On capacity it has far more memory per card, using slower LPDDR.
What is the AI250?
A follow-on with a near-memory computing design that Qualcomm claims has more than 10x the effective memory bandwidth, expected in 2027.
Sources
- Coverage of the October 27, 2025 announcement (ServeTheHome): https://www.servethehome.com/qualcomm-announces-new-integrated-ai-racks-with-768gb-cards-and-a-200mw-ai-deal/
- Coverage of the announcement (DatacenterDynamics): https://www.datacenterdynamics.com/en/news/qualcomm-launches-ai200-and-ai250-chip-offering-targeting-inferencing-workloads-at-rack-scale/
- Qualcomm and Humain release (October 2025): https://www.qualcomm.com/news/releases/2025/10/humain-and-qualcomm-to-deploy-ai-infrastructure-in-saudi-arabia-
- Qualcomm data center roadmap release (June 2026): https://www.qualcomm.com/news/releases/2026/06/qualcomm-unveils-comprehensive-data-center-roadmap-for-the-agent
- Qualcomm and Amazon collaboration release (September 8, 2026): https://investor.qualcomm.com/news-events/press-releases/news-details/2026/Qualcomm-Announces-Multi-Generational-Product-Collaboration-with-Amazon-to-Build-Next-Generation-AI-Data-Center-Infrastructure/default.aspx
- NVIDIA H200: https://www.nvidia.com/en-us/data-center/h200/
- NVIDIA DGX B200: https://www.nvidia.com/en-us/data-center/dgx-b200/