The RTX 4070 Ti Super has 16 GB of GDDR6X on a 256-bit bus, 8,448 CUDA cores and a 285 W total graphics power rating, per NVIDIA's RTX 4070 family page. NVIDIA's launch announcement lists a 672 GB/s memory bandwidth and a $799 price. For AI, 16 GB is the dividing line: it runs 7B to 14B models comfortably and cannot hold a 32B model even at 4-bit.
This is one card in our consumer GPUs for AI guide.
TL;DR
- Specs: 16 GB GDDR6X, 256-bit, 8,448 CUDA cores, 285 W, 672 GB/s.
- Fits (computed): 8B at 8-bit with room to spare, 14B at 4-bit; not 8B at FP16 once overhead is counted, and not 32B at 4-bit.
- Versus 24 GB cards: the RTX 4090 holds a 32B model at 4-bit; this card does not.
- Fine-tuning: QLoRA on 7B to 8B models is the sensible target.
- Use it for: chat, coding assistants and image generation on mid-sized models.
RTX 4070 Ti Super specs
| Spec | RTX 4070 Ti Super | RTX 4090 |
|---|---|---|
| Architecture | Ada Lovelace | Ada Lovelace |
| CUDA cores | 8,448 | 16,384 |
| Boost clock | 2.61 GHz | 2.52 GHz |
| Memory | 16 GB GDDR6X | 24 GB GDDR6X |
| Memory interface | 256-bit | 384-bit |
| Memory bandwidth | 672 GB/s | Not on NVIDIA's page |
| Total graphics power | 285 W | 450 W |
| Launch price | $799, available January 24, 2024 | From $1,599, October 12, 2022 |
| NVLink | Not listed on the product page | No |
The 4070 Ti Super figures are from the RTX 4070 family page and NVIDIA's RTX 40 SUPER announcement (January 8, 2024). The 4090 column is from its product page. Prices are launch list prices.
Note that the 4070 Ti Super has a higher boost clock than the 4090 but about half the cores and two thirds of the memory bus. Per NVIDIA's numbers, memory size and bus width, not clock speed, are what constrain it for LLMs.
What fits in 16 GB
Rule from our VRAM sizing guide: weights are parameters times bytes per parameter, then add the KV cache and about 20% overhead. All numbers below are computed.
| Model | FP16 | 8-bit | 4-bit | Fits in 16 GB? |
|---|---|---|---|---|
| Llama 3.1 8B | 16 GB (19.2 with overhead) | 8 GB (9.6) | 4 GB (4.8) | 8-bit and 4-bit; FP16 does not fit |
| Qwen3-14B (14.8B) | 29.6 GB | 14.8 GB (17.8) | 7.4 GB (8.9) | 4-bit only |
| Qwen3-32B (32.8B) | 65.6 GB | 32.8 GB | 16.4 GB (19.7) | No |
| Llama 3.1 70B | 140 GB | 70 GB | 35 GB | No |
Two practical consequences. An 8B model at 8-bit leaves about 6 GB for the KV cache, which at the Llama 3.1 8B figures in our sizing guide (about 537 MB per 4,096 tokens) is enough for long contexts. And a 14B model at 4-bit leaves roughly 7 GB, comfortable for moderate context but not for 128K.
Use the RTX 4070 Ti VRAM calculator or the RTX 4070 calculator for neighbouring cards; the 4070 Ti Super has its own page at /gpu/rtx-4070-ti-super.
Local LLMs
Ollama is the quickest start:
ollama run gemma4
Pick a model whose quantized file is under about 13 GB so the context has room. llama.cpp lets you offload some layers to system RAM with -ngl when a model slightly exceeds 16 GB, at a large speed cost that depends on your RAM. See what is Ollama and the llama.cpp guide.
Image generation
ComfyUI runs on 16 GB cards. A single image model plus VAE and text encoder is within reach, but stacking several large models or running video models is where 16 GB gets tight, and ComfyUI can offload to system memory when it must. Check each model card for its stated VRAM need. We do not publish an images-per-second figure.
Fine-tuning
Unsloth lists per-model VRAM requirements in its docs. The computed picture: an 8B base in 4-bit is about 4 GB, leaving about 12 GB for LoRA adapters and activations, and a 14B base in 4-bit is about 7.4 GB, leaving much less. Full fine-tuning does not fit. See LoRA fine-tuning and the Unsloth guide.
If a job needs 24 GB or more, rent a card for it instead of changing your model: the RTX 4090 has 24 GB, and the L40S has 48 GB.
Where it sits among neighbours
Within the 16 GB tier, the 4070 Ti Super competes with newer cards that share the 16 GB size. The RTX 5060 Ti guide and the RTX 5070 Ti guide cover the Blackwell cards, and the RTX 5060 Ti vs 5070 Ti comparison lays out their specs. The question to ask of any of them is the same: does the memory size match the largest model you plan to run? Bandwidth then decides how fast tokens come out, because generating a token reads the weights from memory. At 672 GB/s (NVIDIA's figure), the bandwidth ceiling for an 8B model at 8-bit (8 GB of weights) is roughly 672 / 8 = 84 tokens per second, a computed upper bound that ignores the KV cache, overhead and compute, and is not a measurement. Real throughput will be lower.
Power and system fit
At 285 W total graphics power, this is one of the easier cards to fit into an existing PC compared with the 450 W 4090 (NVIDIA's figures). That is a genuine reason to choose it: lower power, lower heat, and a smaller power supply. The trade is the 16 GB ceiling, which is a hard limit, not a tuning problem. If you are unsure which side of 16 GB your workloads fall on, size them first with the VRAM calculator and the sizing guide, and rent a bigger card for the exceptions rather than overbuying for them.
Choosing quantization on 16 GB
On this card the choice is mostly between 8-bit and 4-bit. An 8B model at 8-bit uses about 8 GB of weights and keeps quality close to the original; a 14B model forces 4-bit to fit. Mixed setups work too: run the small model fully on the GPU and offload the large one only when you need it. See the glossary pages on quantization, GGUF and KV cache for what each trade costs.
Run it on a cloud GPU
Run the job that does not fit on 16 GB without buying hardware. The box shows what is available now.
FAQ
How much VRAM does the RTX 4070 Ti Super have?
16 GB of GDDR6X on a 256-bit bus.
Can the RTX 4070 Ti Super run a 32B model?
Not at 4-bit on one card: about 16.4 GB of weights before overhead (computed), more than 16 GB.
Is 16 GB enough for AI?
For 7B to 14B models quantized, yes. For 32B and larger, no.
What is the RTX 4070 Ti Super's memory bandwidth?
672 GB/s, per NVIDIA's January 8, 2024 announcement.
How does it compare with the RTX 4090?
Same architecture, but 16 GB against 24 GB, 8,448 against 16,384 CUDA cores, and a narrower bus. See the 4090 guide.