The NVIDIA RTX PRO 6000 Blackwell is a 96 GB GDDR7 workstation and server GPU with ECC memory, and it is the largest single-card memory pool you can put in a PCIe slot without moving to HBM datacenter parts. For AI work that means one card can hold a 70B model at FP8 or a 32B model at FP16 with room left for the KV cache.
TL;DR
- Memory is the headline. 96 GB of GDDR7 with ECC on every edition (Workstation, Max-Q Workstation, Server), per NVIDIA's product pages.
- Three editions, one memory size. Workstation Edition is 600 W, Max-Q is 300 W, and the Server Edition is passively cooled with a configurable 400 to 600 W power range.
- It fits models a 32 GB RTX 5090 cannot. A 70B model at FP8 is about 70 GB of weights (computed), which only works on a 96 GB card in this class.
- Versus the H100: the H100 SXM has 80 GB of HBM3 at 3.35 TB/s and 900 GB/s NVLink; the PRO 6000 has more memory, lower bandwidth and no NVLink on the specs we could source. Pick by whether your job is capacity-bound or bandwidth-bound.
- Rent before you buy. See the live box below.
What the RTX PRO 6000 Blackwell is
It is the top card of NVIDIA's Blackwell professional line, built for workstations and for rack servers. It uses fifth-generation Tensor Cores and fourth-generation RT Cores, 96 GB of GDDR7 with ECC, and a PCIe Gen 5 x16 interface, according to NVIDIA's RTX PRO 6000 family page and the Workstation Edition page.
Unlike the consumer RTX 5090, it is built for sustained professional use: ECC memory, ISV certification, and (on the Server Edition) MIG partitioning. It sits in the same family as the RTX PRO 5000, which drops to 48 or 72 GB.
The three editions
NVIDIA ships one chip in three forms. All three have 96 GB of GDDR7 with ECC, four DisplayPort 2.1 outputs and PCIe Gen 5 x16, per the family page.
| Edition | Power | Form factor | Cooling |
|---|---|---|---|
| Workstation Edition | 600 W | 5.4 in H x 12 in L, dual slot | Double flow-through |
| Max-Q Workstation Edition | 300 W | 4.4 in H x 10.5 in L, dual slot | Active |
| Server Edition | 400 to 600 W (configurable) | 4.4 in H x 10.5 in L, dual slot (air-cooled); a liquid-cooled single-slot variant also exists | Passive (air-cooled) or liquid |
The Max-Q edition exists so you can fit up to four cards in one chassis within a sane power budget; the family page describes it as suited to dense configurations of up to four GPUs. The Server Edition is the one built for rack deployments and is the version NVIDIA documents with MIG support of up to four fully isolated instances, each with its own memory, cache and compute.
Spec table
Figures below are copied from NVIDIA pages; blanks mean NVIDIA did not publish the number on the page we checked.
| Spec | RTX PRO 6000 Workstation | RTX PRO 6000 Server Edition | H100 SXM |
|---|---|---|---|
| Memory | 96 GB GDDR7, ECC | 96 GB GDDR7, ECC | 80 GB |
| Memory bandwidth | 1,792 GB/s | 1,597 GB/s | 3.35 TB/s |
| CUDA cores | not on page | 24,064 | not compared |
| Power | 600 W | up to 600 W (configurable) | up to 700 W |
| MIG | not on page | up to 4 instances | up to 7 instances at 10 GB |
| Interconnect | PCIe Gen 5 | PCIe Gen 5 | NVLink 900 GB/s |
Sources: Workstation Edition, Server Edition, H100.
On throughput, NVIDIA lists the Workstation Edition at 4,000 AI TOPS (FP4, with sparsity, theoretical peak at boost clock) and 125 TFLOPS of single-precision. The Server Edition page lists 4 PFLOPS at FP4, 2 PFLOPS at FP8 and 1 PFLOPS at FP16/BF16, which are vendor peak figures and not measured results. Those are NVIDIA's numbers and should be read as ceilings.
The two bandwidth figures differ between editions because they come from different NVIDIA pages for different products; check the page for the exact edition you plan to buy or rent.
What fits in 96 GB (computed)
Weights take parameters times bytes per parameter. KV cache and runtime overhead add on top; we use a 15 percent overhead allowance for the "with overhead" column, which is our own rule of thumb and not an NVIDIA figure. See our VRAM sizing guide for the full method, or use the RTX PRO 6000 VRAM calculator.
| Model | Params | FP16 weights | FP8 weights | INT4 weights |
|---|---|---|---|---|
| Llama 3.1 8B | 8B | 16 GB | 8 GB | 4 GB |
| Qwen3-30B-A3B | 30.5B total | 61 GB | 30.5 GB | 15 GB |
| Qwen3-32B | 32.8B | 65.6 GB | 32.8 GB | 16.4 GB |
| Llama 3.1 70B | 70B | 140 GB | 70 GB | 35 GB |
| Qwen3-235B-A22B | 235B total | 470 GB | 235 GB | 117.5 GB |
What that means on 96 GB:
- Qwen3-32B at FP16: 65.6 GB of weights, about 75 GB with the 15 percent allowance. It fits, with roughly 20 GB left for KV cache.
- Llama 3.1 70B at FP8: 70 GB of weights, about 80 GB with the allowance. It fits on one card with around 15 GB for KV cache, which limits long contexts and large batches.
- Llama 3.1 70B at INT4: 35 GB, so it fits with a lot of KV headroom for long context.
- Llama 3.1 70B at FP16: 140 GB. It does not fit on one card; you need two.
- Qwen3-235B-A22B at INT4: 117.5 GB of weights alone. It does not fit on one 96 GB card.
Mixture-of-experts models still need all their weights resident, so Qwen3-30B-A3B is sized by its 30.5B total, not its 3.3B active. See mixture of experts.
Real-world fit
Local and private LLM serving. The sweet spot is a 30 to 70B model served from a single card, because splitting a model across cards over PCIe adds communication cost that a single 96 GB card avoids. Pair it with FP8 or INT4 quantization for 70B-class models.
Fine-tuning. LoRA and QLoRA are the realistic paths. A 4-bit 70B base is about 35 GB of weights (computed), which leaves room for adapters, optimizer state and activations on 96 GB. Full fine-tuning needs roughly 18 bytes per parameter for mixed-precision Adam according to Hugging Face, so even an 8B model wants about 144 GB before activations (computed), which is more than one card holds.
Image and video generation. Large diffusion and video models benefit from the headroom, since the whole pipeline (text encoder, transformer, VAE) can stay resident without offloading.
Multi-tenant serving. On the Server Edition, MIG splits one card into up to four isolated instances, which is useful when many small models share a card.
PRO 6000 versus H100
The H100 SXM has less memory (80 GB) but far more memory bandwidth (3.35 TB/s versus 1,597 to 1,792 GB/s) and 900 GB/s NVLink for multi-GPU scaling. Token generation during decoding is usually limited by memory bandwidth, so a bandwidth gap matters for latency-sensitive serving. The PRO 6000 wins on capacity per card and on being a standard PCIe part. If you want a deeper head-to-head, see the H100 versus RTX PRO 6000 comparison and the RTX PRO 6000 page. For the consumer alternative, read RTX PRO 6000 vs RTX 5090. We do not publish a throughput comparison here because we have no vendor-published like-for-like benchmark to cite.
If you are weighing a small desktop AI box instead, see the DGX Spark guide. For the full buying map across consumer and workstation cards, start at consumer GPUs for AI.
Rent today
Aquanode manages and optimizes GPUs for training and inference workloads. Try the PRO 6000 or an H100 on your own model before committing to hardware.
FAQ
How much VRAM does the RTX PRO 6000 Blackwell have?
96 GB of GDDR7 with ECC, on all three editions, per NVIDIA.
Can it run a 70B model?
Yes at FP8 (about 70 GB of weights, computed) or INT4 (about 35 GB). At FP16 a 70B model needs about 140 GB, so it needs two cards.
What is the difference between Workstation, Max-Q and Server Edition?
Same 96 GB memory. Workstation is 600 W with double flow-through cooling, Max-Q is 300 W and actively cooled for dense multi-card builds, and the Server Edition is passively cooled for racks and lists MIG support.
Does the RTX PRO 6000 support MIG?
NVIDIA's Server Edition page lists up to four fully isolated MIG instances. The family page did not list MIG per edition, so check the datasheet for the edition you buy.
Is it better than an RTX 5090 for AI?
For capacity, yes: 96 GB versus 32 GB. See the side-by-side post.