The AMD Radeon RX 7900 XTX has 24 GB of GDDR6 on a 384-bit bus at up to 960 GB/s, 96 compute units and 96 MB of Infinity Cache, and AMD launched it at a suggested $999. It is the AMD consumer card with the same 24 GB capacity as an RTX 4090, it is on ROCm's supported list, and it runs the common local-inference stacks, but with fewer fast paths than CUDA.
This guide covers specs, what fits in 24 GB, what runs on it today and where it falls short of an NVIDIA card. For all the cards side by side, see Consumer GPUs for AI.
TL;DR
- 24 GB GDDR6, up to 960 GB/s, 355 W typical board power, per AMD's product page.
- It is on AMD's ROCm 7.14.0 supported list (RDNA 3, gfx1100) for Ubuntu 24.04.4, Ubuntu 22.04.5, RHEL 10.1 and RHEL 9.7.
- Ollama lists it for both Linux and Windows. AMD's Windows ROCm story is PyTorch only.
- 24 GB fits 4-bit models up to about 32B, plus a context window. It does not fit a 70B model.
- Verdict: a good 24 GB inference card if you run Linux and llama.cpp, Ollama or vLLM. Choose NVIDIA if your tools assume CUDA, and rent a larger GPU for big models or training.
RX 7900 XTX specs
From AMD's product page and the RDNA 3 launch announcement. Launch price is a suggested US price from late 2022, not a current street price.
| Spec | RX 7900 XTX |
|---|---|
| Architecture | RDNA 3 (chiplet design) |
| Compute units | 96 |
| Stream processors | 6,144 |
| Memory | 24 GB GDDR6, 384-bit, 20 Gbps |
| Memory bandwidth | Up to 960 GB/s |
| Infinity Cache | 96 MB |
| Typical board power | 355 W (800 W PSU recommended) |
| Boost clock | Up to 2,500 MHz |
| Launch price | $999 (AMD, December 2022) |
Sources: AMD RX 7900 XTX page, AMD RDNA 3 announcement.
Against NVIDIA, the capacity matches the RTX 4090 and RTX 3090 at 24 GB. The RTX 5090 has 32 GB. Our RTX 4090 vs RTX 5090 post covers that NVIDIA step.
What fits in 24 GB
This table is computed, not measured: FP16 is 2 bytes per parameter, INT8 is 1, and 4-bit formats are about 0.5 plus overhead. It counts weights only; the KV cache and buffers need another 1 to 4 GB depending on context length.
| Model size | FP16 weights | INT8 weights | 4-bit weights | Fits in 24 GB? |
|---|---|---|---|---|
| 8B | 16 GB | 8 GB | 4 GB | All three, FP16 with a short context |
| 14B | 28 GB | 14 GB | 7 GB | INT8 and 4-bit |
| 32B | 64 GB | 32 GB | 16 GB | 4-bit only |
| 70B | 140 GB | 70 GB | 35 GB | No |
The bandwidth ceiling (bandwidth divided by weight bytes read per token, for a dense model) is 960 GB/s divided by 16 GB, about 60 tokens per second for a 32B model at 4-bit. That is a theoretical upper bound, not a benchmark, and real speeds are lower. For the concepts, see VRAM and quantization, and How much VRAM do I need for LLMs.
What runs on ROCm today
AMD's ROCm 7.14.0 system requirements list the RX 7900 XTX (gfx1100), along with the 7900 XT and 7900 GRE, as supported on Ubuntu 24.04.4, Ubuntu 22.04.5, RHEL 10.1 and RHEL 9.7. Source: ROCm system requirements.
For Radeon on Linux, AMD's documentation (ROCm 7.2.1 edition) lists PyTorch and TensorFlow with training and inference, vLLM, JAX for inference only, llama.cpp, ONNX Runtime with MIGraphX for INT8 and INT4, and FlashAttention-2 with the backward pass. On Windows, PyTorch is the only listed framework. Source: ROCm on Radeon, native Linux compatibility.
llama.cpp and Ollama
llama.cpp supports AMD GPUs through a HIP (ROCm) backend and a Vulkan backend. Commands from llama.cpp's build docs:
# HIP on Windows, with the docs' own example target gfx1100 (the 7900 XTX)
set PATH=%HIP_PATH%\bin;%PATH%
cmake -S . -B build -G Ninja -DGPU_TARGETS=gfx1100 -DGGML_HIP=ON -DCMAKE_C_COMPILER=clang -DCMAKE_CXX_COMPILER=clang++ -DCMAKE_BUILD_TYPE=Release
cmake --build build
# Vulkan on Linux (needs the Vulkan SDK; verify with vulkaninfo)
cmake -B build -DGGML_VULKAN=1
cmake --build build --config Release
On Linux, the HIP build uses hipconfig for HIPCXX and HIP_PATH, as shown in the same document. The 7900 XTX's gfx1100 target matches the Windows example above. Then run a server from a Hugging Face GGUF, per the README:
llama-server -hf ggml-org/Qwen3.5-0.8B-GGUF
Pick a GGUF that fits 24 GB. More flags are in our llama.cpp guide.
Ollama's GPU docs list the RX 7900 series for Linux on ROCm v7, and name the RX 7900 XTX, 7900 XT, 7900 GRE, 7800 XT, 7700 XT, 7600 XT and 7600 for Windows. Vulkan is available for cards outside the ROCm list. Of the AMD consumer cards, this is among the easiest to run in Ollama on Windows. See What is Ollama, and Ollama vs LM Studio for the desktop options.
Honest gaps versus CUDA
- Windows. ROCm on Windows lists only PyTorch. Ollama and llama.cpp work there, but wider tooling is a Linux story.
- OS list. AMD's official list is Ubuntu and RHEL. Other distros work at your own risk.
- Low-precision compute. RDNA 3 launched before the FP8 and FP4 era. We found no AMD statement of native FP8 or FP4 inference support on this card, so 8-bit and 4-bit act as storage formats here. NVIDIA's newer cards add them; see NVFP4 vs MXFP4.
- Training. Many fine-tuning libraries assume CUDA first. Check each project's ROCm notes before you buy for QLoRA or similar work.
- Multi-GPU. No NVLink-style link. Two cards split a model over PCIe.
- Power. 355 W typical board power and an 800 W PSU recommendation.
When to rent instead
If you need more than 24 GB, a model above 32B at 4-bit, or CUDA-only tooling, renting a card for the job is simpler than rebuilding your stack. Aquanode manages and optimizes GPUs for training and inference workloads, and the live box below shows what is available now. The RTX 4090 matches this card's 24 GB, and the RTX 3090 is the older 24 GB NVIDIA option.
Rent today
Compare against NVIDIA cards with the same memory.
FAQ
How much VRAM does the RX 7900 XTX have?
24 GB of GDDR6 on a 384-bit bus, with up to 960 GB/s of bandwidth, per AMD.
Is the RX 7900 XTX supported by ROCm?
Yes. It is on AMD's ROCm 7.14.0 supported list (gfx1100) for Ubuntu and RHEL. On Windows, AMD lists PyTorch as the supported framework.
Can it run a 70B model?
Not entirely in VRAM. A 70B model needs about 35 GB for 4-bit weights (computed). You can offload some layers to system RAM in llama.cpp, with a large speed penalty.
RX 7900 XTX or RTX 4090 for local LLMs?
Both have 24 GB. The 4090 has CUDA support everywhere; the 7900 XTX works with llama.cpp, Ollama and vLLM on ROCm but has a narrower tool list. We do not quote a speed comparison because we did not find an official one.
Does it do image and video generation?
PyTorch runs on it under ROCm, so PyTorch-based diffusion tools can run. Support varies per tool, so check each project's ROCm notes.