For local AI in 2026, the number that decides what a card can do is memory capacity first and memory bandwidth second, not the model name on the box. A 32 GB RTX 5090 is the fastest single consumer card, a 96 GB RTX PRO 6000 is the largest single-GPU workstation card, and unified-memory machines such as DGX Spark and Mac Studio trade speed for capacity.
TL;DR
- Capacity first. If the weights plus KV cache do not fit, speed is irrelevant. Use how much VRAM do I need for LLMs to size the model before you pick a card.
- Best consumer card for AI: the RTX 5090, 32 GB GDDR7 at 1,792 GB/s. The used RTX 3090 and the RTX 4090 share 24 GB.
- More memory than a gaming card: RTX PRO 6000 Blackwell (96 GB) or unified memory (DGX Spark, Mac Studio), at lower bandwidth per gigabyte.
- AMD: the RX 7900 XTX (24 GB) and RX 9070 XT (16 GB) work through ROCm, with a narrower software path than CUDA.
- Renting beats buying when you use the GPU a few days a month, need more memory than any card you can buy, or train. Buying wins for daily, steady use. The break-even section below gives the formula.
The cards, side by side
Memory and bandwidth are vendor-published; where a vendor page omits bandwidth we say so and mark any figure we derived as computed. Each row links our post on the card, and the GPU page where the card exists on Aquanode.
| Card | Memory | Bandwidth | Best for | Read more |
|---|---|---|---|---|
| RTX 5090 | 32 GB GDDR7 | 1,792 GB/s | Fastest single consumer card; 30B-class models at 4-bit, image and video generation | Post, GPU page, vs 4090 |
| RTX 5080 | 16 GB GDDR7 | 960 GB/s | Models up to about 24B at 4-bit, image generation | Post, GPU page |
| RTX 5070 Ti | 16 GB GDDR7 | 896 GB/s | Budget 16 GB Blackwell for small models and diffusion | Post, GPU page |
| RTX 5060 Ti | 16 GB or 8 GB GDDR7 | 448 GB/s | Entry point; buy the 16 GB version | Post, GPU page |
| RTX 4090 | 24 GB GDDR6X | 1,008 GB/s | 24 GB workhorse; QLoRA on 7B to 13B models | Post, GPU page |
| RTX 4070 Ti SUPER | 16 GB GDDR6X | 672 GB/s | Mid-range 16 GB for small models | Post, GPU page |
| RTX 3090 | 24 GB GDDR6X | 936 GB/s | Cheapest way to get 24 GB, used | Post, GPU page |
| RTX PRO 6000 Blackwell | 96 GB GDDR7 ECC | 1,792 GB/s | Largest single-GPU workstation card; 70B at 8-bit | Post, GPU page, vs 5090 |
| RTX PRO 5000 Blackwell | 48 GB GDDR7 ECC (NVIDIA also lists a 72 GB variant) | 1,344 GB/s | 48 GB workstation card | Post, GPU page |
| RTX 6000 Ada | 48 GB GDDR6 ECC | 960 GB/s | Previous-generation 48 GB workstation | Post, GPU page |
| RTX A6000 | 48 GB GDDR6 ECC | 768 GB/s | Older 48 GB, NVLink pair gives 96 GB | Post, GPU page |
| RX 9070 XT | 16 GB GDDR6 | up to 640 GB/s | AMD, ROCm path, smaller models | Post |
| RX 7900 XTX | 24 GB GDDR6 | up to 960 GB/s | AMD with 24 GB, ROCm | Post |
| DGX Spark | 128 GB LPDDR5x unified | 273 GB/s | Large sparse models and prototyping, slow decode | Post, guide |
| Mac Studio | up to 512 GB unified (M5 Ultra) | up to 1.2 TB/s | Big models, no CUDA | Post |
| Jetson Orin Nano | 8 GB LPDDR5 | 102 GB/s | Edge inference at 7 to 25 W | Post |
| Jetson AGX Thor | 128 GB LPDDR5X | 273 GB/s | Robotics, on-device large models | Post |
Sources are in the list at the bottom. The RTX 40 and 30 series pages on NVIDIA's site give memory size and bus width but not bandwidth; the RTX 4090's 1,008 GB/s is from NVIDIA's Ada Lovelace architecture whitepaper, and the RTX 3090's 936 GB/s (19.5 Gbps) is from NVIDIA's GA102 architecture whitepaper. The Mac Studio row shows the top of Apple's range; base configurations carry less memory and bandwidth.
Why bandwidth sets the speed
Generating a token reads the active weights from memory once, so single-stream decode speed is capped at roughly bandwidth divided by model size in bytes. This is computed arithmetic, not a benchmark. A 4-bit 8B model is about 4 GB of weights, so an RTX 5090 has a ceiling near 448 tokens per second (1,792 / 4) and a DGX Spark near 68 (273 / 4); real speeds are lower because of KV cache reads and software overhead. The ratio between the two is what carries over: the 5090 reads memory about 6.6 times faster. Our best GPU for AI guide covers how this plays out up the stack, and dedicated vs shared GPU memory explains why unified-memory machines have lots of capacity and little bandwidth.
Local LLMs
The fit test is weights (parameters times bytes) plus KV cache plus 10 to 20 percent overhead, as worked through in the VRAM guide. By that arithmetic:
| Memory | Largest dense model at 4-bit | At 8-bit |
|---|---|---|
| 16 GB | about 24B | about 12B |
| 24 GB | about 36B | about 18B |
| 32 GB | about 48B | about 24B |
| 48 GB | about 72B | about 36B |
| 96 GB | about 144B | about 72B |
These are computed from bytes per parameter (0.5 at 4-bit, 1 at 8-bit) and reserve a quarter of memory for cache and overhead (weights only get 75 percent), so treat them as approximate. In practice: a 24 GB card runs a 30B-class model at 4-bit with short contexts; a 70B model at 4-bit (35 GB of weights) needs 48 GB or more, or two cards; and a 70B model at 8-bit (70 GB) is what the 96 GB RTX PRO 6000 or a 128 GB unified-memory machine is for.
Mixture-of-experts models change the picture: all weights must be resident, but only a fraction is read per token, which suits high-capacity, low-bandwidth machines such as DGX Spark. Dense models favor high-bandwidth cards. The VRAM calculator has per-card pages, such as RTX 4090 and RTX 5090. Software choices (Ollama, llama.cpp, vLLM) are covered in LLM inference engines.
Image and video generation
Diffusion models are more compute-bound than LLM decoding, so tensor throughput and capacity matter more than bandwidth alone. Newer video models are the capacity pressure: they can want more than 24 GB at full precision, which is where a 32 GB RTX 5090 or a 48 GB to 96 GB workstation card helps. We have not published image-generation benchmarks for these cards and do not repeat third-party numbers here; check each card's post for what the vendor publishes. If you run ComfyUI, persistent cloud GPU for ComfyUI covers keeping your environment.
Fine-tuning at home
Full fine-tuning with AdamW takes about 18 bytes per parameter before activations, per Hugging Face's memory breakdown, which our VRAM guide walks through. That puts full fine-tuning of an 8B model near 144 GB, beyond every consumer card. LoRA and QLoRA are the home-friendly path: freeze a quantized base and train small adapters, so a 24 GB card handles 7B to 13B-class QLoRA by the same arithmetic. See LoRA fine-tuning guide and Unsloth guide for the tools. Training also runs a card at full power for hours: the RTX 5090 is rated at 575 W total graphics power on NVIDIA's page, which is a power-supply and cooling decision, not just a purchase. Multi-GPU at home has a real limit too: consumer cards do not offer the high-bandwidth interconnect of datacenter parts (see what is NVLink).
Datacenter GPUs for comparison
A datacenter card is a different class: the H100 has 80 GB of HBM at 3.35 TB/s, about 1.9 times the bandwidth of an RTX 5090 (3,350 / 1,792, computed). Our datacenter GPUs guide covers the line. For where the two classes overlap, see the RTX PRO 6000 GPU page and the H100 vs RTX 5090 comparison.
When renting beats buying
The break-even is the purchase price divided by the hourly price of the rented GPU you would use instead:
break-even hours = purchase price in dollars / live hourly price in dollars
We do not type an hourly price here: the live box below shows it, and it changes. Per $1 of hourly price, the arithmetic is:
| Item | Price (cited) | Break-even per $1 of hourly price | Continuous days |
|---|---|---|---|
| RTX 5090, launch list | $1,999 (NVIDIA, Jan. 30, 2025) | 1,999 GPU-hours | about 83 |
| RTX 4090, launch list | $1,599 (NVIDIA, Oct. 12, 2022) | 1,599 GPU-hours | about 67 |
| Jetson AGX Thor dev kit | $3,499 (NVIDIA) | 3,499 GPU-hours | about 146 |
| DGX Spark Founders Edition | $4,699 (NVIDIA forum, Feb. 2026) | 4,699 GPU-hours | about 196 |
| Jetson Orin Nano dev kit | $249 (NVIDIA) | 249 GPU-hours | about 10 |
Days are hours divided by 24 and are computed. Read your own number off the box: divide the price by the live hourly figure. Those prices exclude the rest of the computer, power, cooling and your time, and street prices differ from launch lists, so use what you would actually pay. Three things move the answer:
- Utilization. Break-even assumes the card works. A card used a few evenings a month is idle most of the time you own it. Rent, and pay only for hours used.
- Capacity. Past 96 GB per GPU you cannot buy one card at all. Datacenter cards with 80 GB to 288 GB and fast interconnects are rentable by the hour; consumer cards at that scale do not exist.
- Training vs inference. An always-on local assistant favors owning. A weekend fine-tune on a 70B model favors renting a bigger GPU for a few hours. A common split is to own a 24 GB or 32 GB card for daily inference and prototyping, then rent for the heavy runs.
For edge devices, the pattern is explicit: train in the cloud, deploy on the device (Orin Nano, Thor).
Rent today
Aquanode manages and optimizes GPUs for training and inference workloads. The box below shows consumer and workstation cards you can rent on demand, with live prices.
FAQ
What is the best consumer GPU for local LLMs in 2026?
The RTX 5090 for speed and 32 GB. If you need more memory per dollar, a used RTX 3090 gives 24 GB, and workstation cards or unified-memory machines give 48 GB to 512 GB.
Is 16 GB enough for AI?
For 7B to 13B models at 4-bit and for most image generation, yes. For 30B-class models or video generation, 24 GB to 32 GB is more comfortable.
Can AMD cards run AI workloads?
Yes, through ROCm on supported cards, with a narrower software path than CUDA. See our posts on the RX 7900 XTX and RX 9070 XT, and ROCm vs CUDA.
DGX Spark or Mac Studio?
Both give large unified memory with modest bandwidth compared with a discrete GPU. See the comparison.
Is it cheaper to buy a GPU or rent one?
If you will use it most days for months, buying can pay back. Divide the price by the live hourly figure to get your break-even. For occasional or large jobs, renting wins.
Sources
- NVIDIA: compare GeForce graphics cards: memory, bandwidth and TGP for RTX 50 cards, memory sizes for 40 and 30 series
- NVIDIA GeForce RTX 5090: 32 GB GDDR7, 575 W
- NVIDIA GeForce RTX 4090: 24 GB GDDR6X, 384-bit, 450 W
- NVIDIA GeForce RTX 3090: 24 GB GDDR6X
- NVIDIA RTX 50 series launch announcement: RTX 5090 launch list price
- NVIDIA RTX 40 series announcement: RTX 4090 launch price
- NVIDIA RTX PRO 6000 family: 96 GB GDDR7 ECC
- NVIDIA RTX PRO 5000 Blackwell: 48 GB, 1,344 GB/s
- NVIDIA RTX 6000 Ada datasheet and RTX A6000 datasheet: 48 GB, bandwidth
- AMD Radeon RX 9070 XT and RX 7900 XTX: memory and bandwidth
- NVIDIA DGX Spark and price change forum post: 128 GB, 273 GB/s, $4,699
- Apple Mac Studio tech specs: unified memory and bandwidth
- NVIDIA Jetson Orin Nano Super Developer Kit and NVIDIA announcement: 8 GB, 102 GB/s, $249
- NVIDIA Jetson Thor developer blog: 128 GB, 273 GB/s, $3,499
- NVIDIA H100: 80 GB, 3.35 TB/s
- Hugging Face: model memory anatomy: 18 bytes per parameter for mixed-precision AdamW
- NVIDIA Ada Lovelace GPU architecture whitepaper (RTX 4090 memory bandwidth): https://images.nvidia.com/aem-dam/Solutions/geforce/ada/nvidia-ada-gpu-architecture.pdf
- NVIDIA Ampere GA102 GPU architecture whitepaper (RTX 3090 memory bandwidth): https://www.nvidia.com/content/PDF/nvidia-ampere-ga-102-gpu-architecture-whitepaper-v2.pdf