AWS Trainium is Amazon's own AI chip, and Trainium3 is the current generation: AWS lists 2.52 PFLOPs of FP8 and 144 GB of HBM3e per chip, in Trn3 UltraServers of up to 144 chips. It competes with NVIDIA's B200 and H200 on paper, but you can only get it inside AWS, and you program it with the Neuron SDK instead of CUDA.
TL;DR
- Trainium3 is a real peer on specs. 2.52 PFLOPs FP8 and 144 GB of HBM3e at 4.9 TB/s per chip, per AWS.
- It is AWS-only. No other cloud rents it, and no one sells it as a card.
- The software is the trade. Neuron SDK supports PyTorch and JAX, but custom CUDA kernels do not carry over.
- AWS publishes no head-to-head benchmark against NVIDIA. Its performance claims compare Trainium3 with Trainium2.
- Verdict. If you are an AWS shop with a stable, PyTorch-based workload and AWS cost savings matter, Trainium is worth a pilot. If you want portability or the broadest tooling, stay on NVIDIA.
Spec table
Trainium3 figures are from AWS's Trn3 page and launch announcement. HGX B200 figures come from NVIDIA's HGX page, where the FP8 number is quoted with sparsity and the dense figure is half, per NVIDIA's footnote.
| AWS Trainium3 | NVIDIA B200 (HGX B200) | NVIDIA H200 | |
|---|---|---|---|
| FP8 per chip | 2.52 PFLOPs (AWS does not state sparse or dense) | about 4.5 PFLOPS dense (computed: 36 PFLOPS dense per 8-GPU board) | 1,979 TFLOPS dense (3,958 with sparsity) |
| Low-precision formats | MXFP8, MXFP4 | FP8, FP4 | FP8 |
| Memory per chip | 144 GB HBM3e | 1.4 TB across 8 GPUs (about 180 GB each, computed) | 141 GB HBM3e |
| Memory bandwidth | 4.9 TB/s | not listed on the HGX page | 4.8 TB/s |
| Largest scale-up domain | 144 chips (Trn3 UltraServer) | 8 GPUs per HGX board; larger with NVLink rack systems | 8 GPUs per HGX board |
| Scale-up fabric | NeuronSwitch-v1, NeuronLink-v4, 2 TB/s per chip | NVLink 5, 1.8 TB/s per GPU | NVLink 900 GB/s |
| Process | 3nm | not stated on cited page | not stated on cited page |
| Where you rent it | AWS only | Many clouds | Many clouds |
The H200 values come from NVIDIA's H200 page, where the FP8 figure is quoted with sparsity. Read the FP8 rows with care: AWS does not say whether 2.52 PFLOPs assumes sparsity, so do not divide it against NVIDIA's dense number and call it a ratio.
The Trn3 UltraServer
A Trn3 UltraServer holds up to 144 Trainium3 chips for up to 362 FP8 PFLOPs, 20.7 TB of HBM3e and 706 TB/s of aggregate memory bandwidth, per AWS. The chips talk through NeuronSwitch-v1, an all-to-all fabric that AWS says doubles inter-chip bandwidth over Trn2, and AWS cites inter-chip latency "just under 10 microseconds." EC2 UltraClusters 3.0 connect thousands of UltraServers, up to a stated 1 million Trainium chips.
The design parallels NVIDIA's rack-scale systems: one big memory domain per rack, so large models and mixture-of-experts layers do not cross a slow network. For the NVIDIA side of that idea, see our guide to GB200 NVL72.
AWS reports Trn3 went generally available on December 2, 2025.
Generation over generation
AWS's own numbers compare Trainium3 with Trainium2 UltraServers: up to 4.4x higher performance, 3.9x higher memory bandwidth and 4x better performance per watt. On Amazon Bedrock, AWS says Trainium3 is up to 3x faster than Trainium2 with over 5x higher output tokens per megawatt at similar latency per user. These are AWS-versus-AWS figures from AWS's announcement, not independent tests.
One internal inconsistency worth knowing: AWS's launch article lists "4x" energy efficiency in its key takeaways and "40%" better efficiency in the body. We cite only the performance-per-watt line from the AWS announcement page.
Software: Neuron vs CUDA
Trainium runs on the AWS Neuron SDK. AWS lists native PyTorch and JAX support, plus vLLM, Hugging Face Optimum Neuron, PyTorch Lightning and TorchTitan, and says developers can train and deploy "without changing a single line of model code" for supported models. For performance work there is the Neuron Kernel Interface (NKI) for custom kernels and Neuron Explorer for profiling. Trainium also plugs into SageMaker, SageMaker HyperPod, EKS, ECS, AWS Batch and ParallelCluster.
What does not move over:
- Hand-written CUDA kernels. You rewrite them in NKI.
- GPU-only libraries and serving engines without a Neuron backend.
- Anything that assumes NCCL on NVIDIA hardware.
If your model is a standard transformer in PyTorch, the path is short. If your edge is custom kernels, the cost is real.
Who is using it
AWS names Anthropic, Karakuri, Metagenomi, NetoAI, Ricoh and Splash Music as customers reporting training and inference cost reductions of up to 50%, compared with alternatives (AWS does not say specifically against GPUs). Decart reports 4x faster real-time video generation at half the cost of GPUs. Those are customer claims published by AWS.
At the large end, AWS and Anthropic connected more than 500,000 Trainium2 chips in Project Rainier, which AWS calls the largest AI compute cluster in the world and five times the size of the infrastructure used for Anthropic's previous model generation.
Performance: what is published
We could not find an independent head-to-head benchmark of Trainium3 against B200 or H200, and AWS does not publish one. We did not locate Trainium3 entries in the MLPerf rounds we checked, so we make no claim either way; see MLCommons for current submitter lists. Treat any "Trainium beats NVIDIA" headline as unproven until it names the model, precision, batch size and software versions on both sides.
When to choose which
Trainium makes sense if
- Your workload is on AWS already and data egress or latency to other clouds is a concern.
- You run a stable PyTorch or JAX model that Neuron supports.
- You can spend engineering time on a one-off port in exchange for lower unit cost at scale.
- You are comfortable with a single-vendor dependency.
NVIDIA makes sense if
- You want the option to move between clouds, or to your own hardware.
- You depend on CUDA kernels or open-source engines that ship GPU-first.
- You run short experiments where porting time would exceed the savings.
- You need the lowest-risk path for a new model on release day.
Cost
AWS prices Trainium per instance and region, and we do not reproduce those here. To compare fairly, benchmark your own model on both platforms, record tokens per second per chip, and divide by each hourly price. For NVIDIA GPUs on Aquanode, multiply your measured throughput by the live hourly price in the box below.
Rent today
Trainium is rentable only from AWS. If you want NVIDIA hardware by the hour, here is what is live now.
See the B200 and H200 pages for specs and history.
What's next
AWS says Trainium4 is in development, with at least 6x the FP4 processing, 3x the FP8 performance and 4x the memory bandwidth of Trainium3, and that it is being designed to support NVIDIA NVLink Fusion so Trainium and GPU servers can share common MGX racks. These are design targets, not shipping specs, and AWS gives no availability date in the pages we read.
For every datacenter chip on one page, with status, memory and links to each guide, see the datacenter GPU guide.
FAQ
What is AWS Trainium?
Trainium is Amazon's in-house AI accelerator for training and inference, offered through EC2 and services built on it, such as Amazon Bedrock.
Can I buy or rent Trainium outside AWS?
No. It is available through AWS only.
Is Trainium3 faster than the B200?
AWS has not published a comparison. Its claims compare Trainium3 with Trainium2. Peak specs are in a similar range, but speed on your model depends on software and workload.
Do I need to rewrite my code for Trainium?
For supported PyTorch and JAX models, AWS says model code can stay unchanged. Custom CUDA kernels must be rewritten, for example in NKI.
Does Trainium support FP4?
Yes. AWS lists MXFP8 and MXFP4 as supported data types on Trainium3.