RTX PRO 5000 Blackwell: Specs, 48 GB vs 72 GB (2026)

Back
Team Aquanode

Team Aquanode

Sarthak Vaish

Updated OCTOBER 8, 2026Published OCTOBER 8, 2026

The NVIDIA RTX PRO 5000 Blackwell is a 300 W workstation GPU with ECC GDDR7 in two memory sizes, 48 GB and 72 GB, which makes it the middle card of the Blackwell professional line below the RTX PRO 6000. The 72 GB option is what makes it interesting for AI: it holds models that a 48 GB card cannot.

TL;DR

  • Two memory options. 48 GB or 72 GB of GDDR7 with ECC, at 1,344 GB/s of bandwidth, per NVIDIA's product page.
  • 300 W, dual slot, active cooling, PCIe Gen 5. It runs in a normal workstation without a power-supply upgrade.
  • MIG support: up to two fully isolated instances, each with its own memory, cache and compute.
  • Best fit: 30B to 32B models at FP8, 70B models at INT4, and LoRA fine-tuning of mid-size models. A 70B model at FP8 (70 GB of weights, computed) is too tight even on the 72 GB version.
  • Not a benchmark claim. NVIDIA's peak AI figure is 2,064 AI TOPS; we do not turn that into a speed estimate.

What the RTX PRO 5000 is

It is a Blackwell professional card with fifth-generation Tensor Cores, fourth-generation RT Cores, three NVENC and three NVDEC engines, and four DisplayPort 2.1 outputs, according to NVIDIA's RTX PRO 5000 page. The form factor is dual slot at 4.4 in by 10.5 in. NVIDIA did not list a CUDA core count on that page, so we do not state one.

You can look at the card on its GPU page and size workloads with the RTX PRO 5000 VRAM calculator.

Spec table against its neighbours

SpecRTX PRO 5000RTX PRO 6000 WorkstationRTX PRO 6000 Max-Q
Memory48 GB or 72 GB GDDR7, ECC96 GB GDDR7, ECC96 GB GDDR7, ECC
Memory bandwidth1,344 GB/s1,792 GB/snot on page
Power300 W600 W300 W
Form factordual slot, 4.4 in x 10.5 indual slot, 5.4 in x 12 indual slot, 4.4 in x 10.5 in
Coolingactivedouble flow-throughactive
AI performance2,064 AI TOPS4,000 AI TOPS (FP4, sparsity)not on page
MIGup to 2 instancessee Server Edition pagenot listed
PCIeGen 5Gen 5Gen 5

Sources: PRO 5000, PRO 6000 Workstation Edition, PRO 6000 family.

The PRO 5000 and the PRO 6000 Max-Q draw the same 300 W in the same chassis size. The PRO 6000 gives you 96 GB where the PRO 5000 gives you 48 or 72 GB, so the choice is mostly about how much memory you need per slot. NVIDIA's AI TOPS entries come from different product pages and may use different precision and sparsity assumptions, so we show them for reference only.

What fits in 48 GB and 72 GB (computed)

Weights are parameters times bytes per parameter. The allowance column adds 15 percent for runtime overhead, our own rule of thumb; KV cache comes on top. Full method in our VRAM sizing guide.

ModelFP16 weightsFP8 weightsINT4 weights
Llama 3.1 8B (8B)16 GB8 GB4 GB
Qwen3-30B-A3B (30.5B total)61 GB30.5 GB15 GB
Qwen3-32B (32.8B)65.6 GB32.8 GB16.4 GB
Llama 3.1 70B (70B)140 GB70 GB35 GB
Qwen3-235B-A22B (235B total)470 GB235 GB117.5 GB

Reading the table against the two memory sizes:

  • 48 GB version. Llama 3.1 8B at FP16 (18.4 GB with allowance), Qwen3-32B at FP8 (37.7 GB with allowance), and Llama 3.1 70B at INT4 (40.3 GB with allowance, leaving under 8 GB for KV cache) all fit. A 70B model at INT4 on 48 GB works for short contexts, not for long ones.
  • 72 GB version. Qwen3-30B-A3B at FP16 needs 61 GB of weights and about 70 GB with the allowance, so it only just fits. Qwen3-32B at FP16 needs 65.6 GB of weights and about 75 GB with the allowance, which exceeds 72 GB. Use FP8 for that model. Llama 3.1 70B at INT4 has about 30 GB of headroom for KV cache.
  • Neither version. Llama 3.1 70B at FP8 is 70 GB of weights, about 80 GB with the allowance; it does not fit on 72 GB. That needs the 96 GB PRO 6000.
  • Beyond one card. Qwen3-235B-A22B at INT4 is 117.5 GB of weights, which is more than any single PRO 5000.

Mixture-of-experts models are sized by total parameters, not active ones; see mixture of experts.

Real-world fit

Local LLM serving. The 72 GB version is a good single-card home for 30B-class models at FP8 and 70B-class models at 4-bit. Pick FP8 or quantization to match the table above.

Fine-tuning. LoRA and QLoRA on 7B to 32B models fit comfortably. Full fine-tuning with mixed-precision AdamW needs about 18 bytes per parameter according to Hugging Face, so a 3B model is already about 54 GB before activations (computed). QLoRA on a 70B base needs 35 GB of 4-bit weights plus adapters, optimizer state and activations, which is realistic on the 72 GB version and tight on 48 GB.

Image and video generation. Most image models fit with room to spare. Video models and large batch sizes benefit from the 72 GB option.

Sharing a card. NVIDIA lists MIG with up to two isolated instances, which suits a small team splitting one card into two jobs with guaranteed quality of service.

PRO 5000 versus the alternatives

Against the RTX PRO 6000, you give up memory (up to 72 GB versus 96 GB) and bandwidth (1,344 versus 1,792 GB/s for the Workstation Edition), and you gain a 300 W power budget that matches the Max-Q. Against the older 48 GB cards, the RTX 6000 Ada and RTX A6000, you get GDDR7, more bandwidth than either (960 and 768 GB/s respectively) and an optional 72 GB size. Against consumer cards, the PRO 5000 offers ECC and more memory per slot than a 32 GB RTX 5090. For the full map read consumer GPUs for AI. We do not quote a throughput comparison because we have no vendor-published like-for-like benchmark.

Rent today

Aquanode manages and optimizes GPUs for training and inference workloads. Try a PRO 6000 class card on your model and see how much memory you really need before choosing between 48, 72 and 96 GB.

FAQ

How much VRAM does the RTX PRO 5000 Blackwell have?

NVIDIA lists 48 GB or 72 GB of GDDR7 with ECC, depending on the configuration.

What is the power draw?

NVIDIA lists a 300 W maximum, dual slot with active cooling.

Can it run a 70B model?

At INT4 yes (35 GB of weights, computed). At FP8 the weights alone are 70 GB, which does not fit on the 72 GB version once you add overhead and KV cache.

Does it support MIG?

Yes, NVIDIA lists up to two fully isolated instances, each with its own memory, cache and compute.

Is it better than an RTX 6000 Ada?

It has the newer architecture, GDDR7 and higher bandwidth (1,344 versus 960 GB/s), and an optional 72 GB size. See the RTX 6000 Ada guide.

Sources

#consumer gpu#workstation gpu#rtx pro 5000#blackwell#vram#llm inference

Submit the job. Everything after that is ours.

Sign up in 60 seconds. Pay for the GPU minutes you actually use.

© 2026 Aquanode. All rights reserved.

All trademarks, logos and brand names are the property of their respective owners.