The NVIDIA RTX 6000 Ada Generation is a 48 GB GDDR6 workstation GPU with ECC, 18,176 CUDA cores and a 300 W power limit. For AI it is a dual-slot card that holds a 70B model at 4-bit or a 32B model at FP8, and it is the direct predecessor of the Blackwell RTX PRO 6000 and RTX PRO 5000.
TL;DR
- 48 GB GDDR6 with ECC on a 384-bit bus at 960 GB/s, per NVIDIA's datasheet.
- 18,176 CUDA cores, 568 fourth-generation Tensor Cores, 91.1 TFLOPS single precision (vendor peak).
- 300 W, dual slot, PCIe Gen 4 x16. It fits ordinary workstations.
- Good for: 32B models at FP8, 70B models at INT4 with short context, LoRA and QLoRA fine-tuning of 7B to 32B models.
- Versus the RTX A6000: same 48 GB, but a newer architecture and 960 versus 768 GB/s of bandwidth. See the RTX A6000 guide.
What it is
The RTX 6000 Ada is NVIDIA's Ada Lovelace workstation flagship, the generation after the Ampere RTX A6000 and before Blackwell. The datasheet lists 142 third-generation RT Cores, 568 fourth-generation Tensor Cores and 18,176 CUDA cores with 48 GB of ECC graphics memory. NVIDIA's product page lists a 300 W maximum, a dual-slot 4.4 in by 10.5 in form factor with active cooling, four DisplayPort 1.4 outputs and a PCIe Gen 4 x16 bus, and gives 1,457 AI TOPS as theoretical FP8 with sparsity.
The GPU page is here and the sizing calculator is here.
Spec table
| Spec | RTX 6000 Ada | RTX A6000 | RTX PRO 5000 |
|---|---|---|---|
| Architecture | Ada Lovelace | Ampere | Blackwell |
| Memory | 48 GB GDDR6, ECC | 48 GB GDDR6, ECC | 48 or 72 GB GDDR7, ECC |
| Memory bandwidth | 960 GB/s | 768 GB/s | 1,344 GB/s |
| CUDA cores | 18,176 | 10,752 | not on page |
| Tensor cores | 568 (4th gen) | 336 (3rd gen) | 5th gen |
| FP32 peak | 91.1 TFLOPS | 38.7 TFLOPS | not on page |
| Power | 300 W | 300 W | 300 W |
| PCIe | Gen 4 | Gen 4 | Gen 5 |
| NVLink | not listed | 112.5 GB/s bridge | not listed |
Sources: RTX 6000 Ada datasheet, RTX A6000 datasheet, RTX PRO 5000. All three draw 300 W, so the real differences are architecture, bandwidth and memory type.
The two older cards both stop at 48 GB, which is the main limit of this generation. The Ada card is the better of the two for bandwidth-bound work such as token generation, because bandwidth is 960 versus 768 GB/s (a 25 percent gap, computed).
What fits in 48 GB (computed)
Weights are parameters times bytes per parameter. The allowance adds 15 percent for runtime overhead, our own rule of thumb, and KV cache comes on top. See the VRAM sizing guide.
| Model | FP16 weights | FP8 weights | INT4 weights |
|---|---|---|---|
| Llama 3.1 8B | 16 GB | 8 GB | 4 GB |
| Qwen3-30B-A3B (30.5B total) | 61 GB | 30.5 GB | 15 GB |
| Qwen3-32B (32.8B) | 65.6 GB | 32.8 GB | 16.4 GB |
| Llama 3.1 70B | 140 GB | 70 GB | 35 GB |
- Fits: Llama 3.1 8B at FP16 (18.4 GB with allowance), Qwen3-30B-A3B at FP8 (35 GB with allowance), Qwen3-32B at FP8 (37.7 GB with allowance, about 10 GB left for KV cache), and Llama 3.1 70B at INT4 (40.3 GB with allowance, under 8 GB left for KV cache).
- Does not fit: Qwen3-32B at FP16 (65.6 GB of weights), Qwen3-30B-A3B at FP16 (61 GB), and Llama 3.1 70B at FP8 (70 GB).
Note that this card has no FP4 Tensor Core path; FP4 hardware arrived with Blackwell, per NVIDIA's Blackwell PRO 6000 pages. INT4 weight-only quantization such as GPTQ or AWQ still saves memory on Ada; it dequantizes to higher precision for compute. See quantization.
Real-world fit
Local LLM serving. 32B models at FP8 and 70B at 4-bit run on one card. Long contexts on a 70B model are limited by the leftover KV budget.
Fine-tuning. LoRA and QLoRA on 7B to 32B models fit in 48 GB. Full fine-tuning at about 18 bytes per parameter (Hugging Face) puts a 2B model at 36 GB before activations (computed), so full fine-tuning past a couple of billion parameters needs more memory.
Image and video. 48 GB is generous for image models and workable for many video models.
Multi-card. NVIDIA's datasheet and product page do not list NVLink for this card, so multi-card setups talk over PCIe. The older RTX A6000 does list a 112.5 GB/s NVLink bridge between two cards.
Which to choose
If you want the cheapest 48 GB with modern Tensor Cores, the 6000 Ada fits. If you need more than 48 GB per card, step to the Blackwell RTX PRO 5000 (72 GB) or RTX PRO 6000 (96 GB). If you need datacenter-style serving, compare it with an L40S, which also has 48 GB. See the RTX 6000 Ada versus RTX A6000 comparison for a spec table, and consumer GPUs for AI for the wider buying map. We do not publish a speed ratio because we have no vendor-published like-for-like benchmark.
How much context is left for the KV cache (computed)
Weights are only part of the budget. The KV cache formula from NVIDIA's inference optimization guide is 2 x layers x KV heads x head dimension x sequence length x batch x bytes. Using the Llama 3.1 70B configuration (80 layers, 8 KV heads, head dimension 128, per the model's config.json on Hugging Face, via the Unsloth mirror: 80 hidden layers, 64 attention heads, 8 key-value heads, hidden size 8,192, so head dimension 8,192 / 64 = 128) at FP16 cache precision, that is 2 x 80 x 8 x 128 x 2 bytes, about 0.33 MB per token (computed). An 8,192-token context is then about 2.7 GB and a 32,768-token context about 10.7 GB, per sequence.
On a 48 GB card running the 70B model at INT4, the weights plus our 15 percent allowance take about 40.3 GB, leaving roughly 7.7 GB. That is around 23,000 tokens of KV cache in total across all concurrent sequences (computed), so a handful of 4,096-token chats fit and a single 32K-token context does not. If you need long contexts or many users on a 70B model, you need more memory per card or more cards. Quantizing the KV cache or using a smaller model are the other levers; see KV cache.
Check what you have
Before loading anything, confirm the card and its memory from the driver. This standard nvidia-smi query is documented in NVIDIA's nvidia-smi manual:
nvidia-smi --query-gpu=name,memory.total,memory.used --format=csv
It prints the board name and total and used memory, so you can compare against the weight tables above before a model fails to load.
Rent today
Aquanode manages and optimizes GPUs for training and inference workloads. Test your model on a 48 GB class card before deciding whether you need more memory.
FAQ
How much VRAM does the RTX 6000 Ada have?
48 GB of GDDR6 with ECC.
What is the memory bandwidth?
960 GB/s on a 384-bit interface, per NVIDIA's datasheet.
Can it run Llama 3.1 70B?
At INT4, yes: 35 GB of weights (computed), about 40 GB with overhead. At FP8 it needs 70 GB and does not fit.
Is the RTX 6000 Ada the same as the RTX A6000?
No. The RTX A6000 is the Ampere generation with 768 GB/s and 10,752 CUDA cores. The 6000 Ada has 960 GB/s and 18,176 CUDA cores.
Does it support NVLink?
NVIDIA's pages for this card do not list NVLink.