The AMD Radeon RX 9070 XT has 16 GB of GDDR6 on a 256-bit bus (up to 640 GB/s), 64 compute units and 128 AI accelerators, and AMD launched it at a suggested price of $599. For AI it is a capable 16 GB card for models up to roughly 14B parameters at 4-bit, and it is officially supported on ROCm, but the software path is narrower than on an NVIDIA card of the same size.
This guide covers the specs, what fits in 16 GB, what runs on the card today (ROCm, llama.cpp, Ollama), and when to rent a bigger GPU instead. For the full map of cards, see Consumer GPUs for AI.
TL;DR
- 16 GB of GDDR6, up to 640 GB/s, 304 W typical board power, per AMD's product page.
- The RX 9070 XT (RDNA 4, gfx1201) is on AMD's ROCm Linux support list for ROCm 7.14.0, on Ubuntu 24.04.4, Ubuntu 22.04.5, RHEL 10.1 and RHEL 9.7.
- Ollama and llama.cpp both run on it. Ollama lists the 9070 series for Linux on ROCm v7; the Windows list in Ollama's docs names the 7000 series, with Vulkan as an extra option.
- Gap versus CUDA: fewer prebuilt packages, a Linux-first story, and fine-tuning tools that assume CUDA first. If your stack is PyTorch, vLLM or llama.cpp on Linux, it works. If you depend on a niche CUDA library, check first.
- Verdict: a good value card for local inference on 7B to 14B models. For anything larger, or for training, rent a bigger GPU by the hour.
RX 9070 XT specs
Figures are from AMD's product page and launch press release unless noted. Launch price is a US suggested price from early 2025, not a current street price.
| Spec | RX 9070 XT |
|---|---|
| Architecture | RDNA 4 |
| Compute units | 64 |
| AI accelerators | 128 |
| Stream processors | 4,096 |
| Memory | 16 GB GDDR6, 256-bit, up to 20 Gbps |
| Memory bandwidth | Up to 640 GB/s |
| Infinity Cache | 64 MB |
| Typical board power | 304 W (750 W PSU recommended) |
| Peak FP32 | 48.7 TFLOPS (AMD's figure) |
| Launch price | $599 (AMD, announced February 28, 2025) |
Sources: AMD RX 9070 XT page, AMD launch press release. We did not find a published low-precision (FP8 or INT8) throughput figure on the product page, so we do not quote one.
The 16 GB figure is the one that matters. For LLM inference the limit is almost always memory capacity first and memory bandwidth second, and 16 GB puts the card in the same capacity class as the RTX 5070 Ti and RTX 5060 Ti 16 GB. The RTX 4090 has 24 GB, and the RTX 5090 has 32 GB. See How much VRAM do I need for LLMs for the general rule.
What fits in 16 GB
Weights take roughly parameters times bytes per parameter. This table is computed, not measured: FP16 is 2 bytes, INT8 is 1 byte, and 4-bit formats (GGUF Q4 and similar) are about 0.5 bytes plus a little overhead. It counts weights only. The KV cache for your context length and runtime buffers come on top, so leave 1 to 3 GB of headroom.
| Model size | FP16 weights | INT8 weights | 4-bit weights | Fits in 16 GB? |
|---|---|---|---|---|
| 7B to 8B | 14 to 16 GB | 7 to 8 GB | 3.5 to 4 GB | INT8 and 4-bit yes; FP16 is too tight |
| 14B | 28 GB | 14 GB | 7 GB | 4-bit yes; INT8 barely |
| 24B | 48 GB | 24 GB | 12 GB | 4-bit yes with a short context |
| 32B | 64 GB | 32 GB | 16 GB | No (weights alone fill the card) |
| 70B | 140 GB | 70 GB | 35 GB | No |
A bandwidth ceiling is also computable. Generating one token reads roughly all the weights once for a dense model, so tokens per second cannot exceed bandwidth divided by weight size. At 640 GB/s, an 8 GB model (8B at INT8) tops out near 80 tokens per second, and a 12 GB model (24B at 4-bit) near 53. These are theoretical upper bounds that real runs fall below, not benchmarks. Use our glossary on VRAM and quantization for the underlying ideas.
What runs on ROCm today
AMD's ROCm system requirements page (ROCm 7.14.0) marks the RX 9070 XT, RX 9070 and RX 9070 GRE as supported, all RDNA 4 with target gfx1201, on Ubuntu 24.04.4, Ubuntu 22.04.5, RHEL 10.1 and RHEL 9.7. The RX 9060 XT family (gfx1200) is also listed. Source: ROCm system requirements.
AMD's Radeon documentation (ROCm 7.2.1 edition) lists, for Radeon on Linux:
- PyTorch and TensorFlow with training and inference support.
- vLLM, listed as supported.
- JAX, inference only.
- llama.cpp for efficient inference.
- ONNX Runtime with MIGraphX for INT8 and INT4 inference.
- FlashAttention-2 with the backward pass enabled.
On Windows, AMD's Radeon page lists PyTorch as the only framework, covering the 9000 series and select 7000 series cards. Source: ROCm on Radeon, native Linux compatibility. That page did not name individual cards in the text we could read, so use the system requirements page above for the per-card answer.
llama.cpp and Ollama on the 9070 XT
llama.cpp has two paths for AMD GPUs. The HIP backend uses ROCm. The Vulkan backend works through the graphics driver and needs no ROCm install. From llama.cpp's own build docs:
# Vulkan backend (needs the Vulkan SDK; verify with vulkaninfo)
cmake -B build -DGGML_VULKAN=1
cmake --build build --config Release
# HIP / ROCm backend on Linux (the docs show gfx1030 as the example target)
HIPCXX="$(hipconfig -l)/clang" HIP_PATH="$(hipconfig -R)" \
cmake -S . -B build -DGGML_HIP=ON -DGPU_TARGETS=gfx1030 -DCMAKE_BUILD_TYPE=Release \
&& cmake --build build --config Release -- -j 16
For the 9070 XT, set GPU_TARGETS=gfx1201 (the target AMD lists above). The llama.cpp page does not list gfx1201 in its example, and omitting GPU_TARGETS builds for the GPUs in your system. Then serve a model straight from Hugging Face, using the command in the llama.cpp README:
llama-server -hf ggml-org/Qwen3.5-0.8B-GGUF
Swap in a GGUF that fits your 16 GB. See the llama.cpp guide for flags and multi-GPU options.
Ollama's GPU docs say Linux needs the AMD ROCm v7 driver stack and list the RX 9070 series among supported cards. On Windows its list names the 7900 XTX, 7900 XT, 7900 GRE, 7800 XT, 7700 XT, 7600 XT and 7600, and Vulkan is available as a backend for cards outside the ROCm list, on by default and disabled with OLLAMA_VULKAN=0. In short, a 9070 XT user on Linux gets the ROCm path, and on Windows should expect Vulkan. For the tool itself, see What is Ollama.
Honest gaps versus CUDA
- Packaging. The supported-OS list is short (Ubuntu and RHEL). On other distros you are on your own.
- Windows. AMD lists only PyTorch there. Linux is where the full list lives.
- FP8 and FP4. NVIDIA's recent cards add native FP4 paths used by NVFP4. We found no AMD statement of equivalent FP4 inference support on this card, so treat 4-bit as a weight-storage format, not a compute format.
- Training tools. Fine-tuning libraries are written CUDA first. Many work on ROCm, and some do not. Check a library's own docs for ROCm before you buy for training, and note 16 GB limits you to small models with QLoRA either way.
- Multi-GPU. There is no NVLink-style link on this card.
None of this makes the card a bad buy for inference. It makes it a card where you verify your stack first.
When to rent a GPU instead
If your model is above about 14B parameters at 4-bit, if you need 32 GB or more of VRAM, or if you want CUDA-only tooling for a one-off run, renting is simpler than buying. An RTX 4090 has 24 GB, and an RTX 5070 Ti is a 16 GB NVIDIA comparison point. Aquanode manages and optimizes GPUs for training and inference workloads, and the live box below shows what is available now.
Rent today
Try a model on a CUDA card before you decide what to buy.
FAQ
How much VRAM does the RX 9070 XT have?
16 GB of GDDR6 on a 256-bit bus, with up to 640 GB/s of bandwidth, per AMD's product page.
Is the RX 9070 XT supported by ROCm?
Yes. AMD's ROCm 7.14.0 system requirements list it as supported (gfx1201) on Ubuntu 24.04.4, Ubuntu 22.04.5, RHEL 10.1 and RHEL 9.7.
Can the RX 9070 XT run a 70B model?
Not on its own. A 70B model needs about 35 GB for 4-bit weights (computed), more than twice the card's 16 GB. You can offload layers to system RAM in llama.cpp, but speed drops sharply.
Should I pick Vulkan or ROCm in llama.cpp?
Vulkan is the easier install and works through the graphics driver. HIP uses ROCm and is the route AMD documents for compute. Try Vulkan first, then HIP if you want to compare. We did not find an official speed comparison to cite.
RX 9070 XT or RX 7900 XTX for AI?
The 7900 XTX has 24 GB and 960 GB/s, which fits larger models. See the RX 7900 XTX guide.