RTX A6000: Specs, 48 GB VRAM, NVLink, AI Guide (2026)

Back
Team Aquanode

Team Aquanode

Sarthak Vaish

Updated OCTOBER 8, 2026Published OCTOBER 8, 2026

The NVIDIA RTX A6000 is an Ampere workstation GPU with 48 GB of GDDR6 ECC memory, 10,752 CUDA cores and a 300 W power limit, and it is the one older workstation card that can pair with a second card over NVLink for a combined 96 GB. In 2026 it is still a practical 48 GB card for 4-bit 70B models and mid-size fine-tuning.

TL;DR

  • 48 GB GDDR6 with ECC, 384-bit bus, 768 GB/s of bandwidth, per NVIDIA's datasheet.
  • 10,752 CUDA cores, 336 third-generation Tensor Cores, 38.7 TFLOPS single-precision peak.
  • NVLink: a bridge connects two cards at 112.5 GB/s (bidirectional) for a combined 96 GB, per NVIDIA.
  • No FP8 or FP4 Tensor Core hardware. Those arrived with Ada (FP8) and Blackwell (FP4). Weight-only INT4 still saves memory.
  • Fits: 70B at INT4 (short context), 32B at INT8 or 8-bit, 8B at FP16 with plenty of room.

What it is

The RTX A6000 is the Ampere-generation predecessor of the RTX 6000 Ada. NVIDIA's datasheet lists 48 GB GDDR6 on a 384-bit interface, 768 GB/s of memory bandwidth, 10,752 CUDA cores, 336 third-generation Tensor Cores, 38.7 TFLOPS single precision, 75.6 TFLOPS RT performance and 309.7 TFLOPS tensor performance (listed with sparsity). Total board power is 300 W through one 8-pin connector. The product page lists a dual-slot 4.4 in by 10.5 in card with active cooling, PCIe Gen 4 x16 and four DisplayPort 1.4a outputs.

Details for this card are on its GPU page and in the RTX A6000 VRAM calculator.

Spec table

SpecRTX A6000RTX 6000 AdaRTX PRO 6000 Workstation
ArchitectureAmpereAda LovelaceBlackwell
Memory48 GB GDDR6, ECC48 GB GDDR6, ECC96 GB GDDR7, ECC
Bandwidth768 GB/s960 GB/s1,792 GB/s
CUDA cores10,75218,176not on page
Power300 W300 W600 W
PCIeGen 4Gen 4Gen 5
NVLink112.5 GB/s bridge, 2 cardsnot listednot listed

Sources: A6000 datasheet, RTX 6000 Ada datasheet, PRO 6000. A direct spec table lives on the RTX 6000 Ada versus RTX A6000 comparison.

The 112.5 GB/s NVLink figure is the bridge between two A6000 cards. It is much lower than datacenter NVLink, so a two-card A6000 setup shares memory but does not behave like one large card, and software still has to shard the model with tensor parallelism.

What fits in 48 GB (computed)

Weights are parameters times bytes per parameter. The allowance adds 15 percent for runtime overhead, our own rule of thumb; KV cache comes on top. Method in the VRAM sizing guide.

ModelFP16 weights8-bit weightsINT4 weights
Llama 3.1 8B16 GB8 GB4 GB
Qwen3-30B-A3B (30.5B total)61 GB30.5 GB15 GB
Qwen3-32B (32.8B)65.6 GB32.8 GB16.4 GB
Llama 3.1 70B140 GB70 GB35 GB
  • One card (48 GB): Llama 3.1 8B at FP16 (18.4 GB with allowance), Qwen3-32B at 8-bit (37.7 GB with allowance), Qwen3-30B-A3B at 8-bit (35 GB with allowance), and Llama 3.1 70B at INT4 (40.3 GB with allowance, under 8 GB left for KV cache).
  • Two cards over NVLink (96 GB combined): Qwen3-32B at FP16 (65.6 GB of weights, about 75 GB with allowance) and Llama 3.1 70B at 8-bit (70 GB of weights, about 80 GB with allowance). Each card holds half, so the 48 GB per-card limit still applies to each shard.
  • Does not fit even on two cards: Llama 3.1 70B at FP16 (140 GB).

On this card 8-bit means INT8 or weight-only 8-bit, not the hardware FP8 path found on newer GPUs. Weight-only formats such as GPTQ and AWQ save memory but dequantize for compute. See quantization.

Real-world fit

Local LLM serving. The A6000 handles 7B to 32B models comfortably and 70B at 4-bit with a short context. Token generation is limited by the 768 GB/s of memory bandwidth, which is the lowest of the three 48 GB-class cards compared above.

Fine-tuning. LoRA and QLoRA for 7B to 32B models fit in 48 GB. At about 18 bytes per parameter for full mixed-precision Adam (Hugging Face), full fine-tuning tops out around a 2B model on one card (36 GB, computed, before activations).

Image and video generation. 48 GB is plenty for common image models, and the ECC memory suits long jobs.

Why people still buy it. It is a 48 GB card with ECC and NVLink on the used and refurbished market. We do not quote used prices because they vary by seller and day. For the newer 48 GB option, see the RTX 6000 Ada guide. For more memory, see the RTX PRO 6000 Blackwell guide.

A6000 versus newer options

The 6000 Ada has 960 GB/s (25 percent more bandwidth, computed) and 18,176 CUDA cores. The Blackwell PRO 6000 doubles the memory to 96 GB and adds FP4 hardware. The consumer RTX 3090 also pairs two cards over NVLink, but with 24 GB each and no ECC. For the wider picture of consumer and workstation options, read consumer GPUs for AI. We do not publish a throughput comparison because we have no vendor-published like-for-like benchmark.

How much context is left for the KV cache (computed)

Weights are only part of the budget. The KV cache formula from NVIDIA's inference optimization guide is 2 x layers x KV heads x head dimension x sequence length x batch x bytes. Using the Llama 3.1 70B configuration (80 layers, 8 KV heads, head dimension 128, per the model's config.json on Hugging Face, via the Unsloth mirror: 80 hidden layers, 64 attention heads, 8 key-value heads, hidden size 8,192, so head dimension 8,192 / 64 = 128) at FP16 cache precision, that is 2 x 80 x 8 x 128 x 2 bytes, about 0.33 MB per token (computed). An 8,192-token context is then about 2.7 GB and a 32,768-token context about 10.7 GB, per sequence.

On a 48 GB card running the 70B model at INT4, the weights plus our 15 percent allowance take about 40.3 GB, leaving roughly 7.7 GB. That is around 23,000 tokens of KV cache in total across all concurrent sequences (computed), so a handful of 4,096-token chats fit and a single 32K-token context does not. If you need long contexts or many users on a 70B model, you need more memory per card or more cards. Quantizing the KV cache or using a smaller model are the other levers; see KV cache.

Check what you have

Before loading anything, confirm the card and its memory from the driver. This standard nvidia-smi query is documented in NVIDIA's nvidia-smi manual:

nvidia-smi --query-gpu=name,memory.total,memory.used --format=csv

It prints the board name and total and used memory, so you can compare against the weight tables above before a model fails to load.

Rent today

Aquanode manages and optimizes GPUs for training and inference workloads. Check whether 48 GB is enough for your model before you buy.

FAQ

How much VRAM does the RTX A6000 have?

48 GB of GDDR6 with ECC, on a 384-bit bus at 768 GB/s.

Does the RTX A6000 support NVLink?

Yes. NVIDIA lists a bridge connecting two cards with 112.5 GB/s (bidirectional) and a combined 96 GB.

Can it run Llama 3.1 70B?

At INT4 on one card: 35 GB of weights (computed), about 40 GB with overhead. At 8-bit it needs two cards.

Is the RTX A6000 good for AI in 2026?

It is capable for 48 GB inference and fine-tuning, but it lacks FP8 and FP4 hardware and has the lowest bandwidth of the 48 GB workstation cards covered here.

What replaced the RTX A6000?

The RTX 6000 Ada Generation, which keeps 48 GB but raises bandwidth to 960 GB/s. See the RTX 6000 Ada guide.

Sources

#consumer gpu#workstation gpu#rtx a6000#ampere#vram#llm inference

Submit the job. Everything after that is ours.

Sign up in 60 seconds. Pay for the GPU minutes you actually use.

© 2026 Aquanode. All rights reserved.

All trademarks, logos and brand names are the property of their respective owners.