RTX 3070 Ti VRAM: 8 GB GDDR6X, and what fits in it
Ampere. 8 GB VRAM, verified from the vendor datasheet. No current live pricing to show.
8 GB
VRAM
Ampere
Architecture
–
From
GDDR6X
Memory
Models that fit at FP16 (8 GB)
- google-bert/bert-base-uncased (110M)
- Qwen/Qwen3-0.6B (752M)
- openai-community/gpt2 (137M)
- Qwen/Qwen2.5-1.5B-Instruct (1.5B)
- Qwen/Qwen2.5-3B-Instruct (3.1B)
- openai/whisper-large-v3-turbo (809M)
- Qwen/Qwen3-Embedding-0.6B (596M)
- meta-llama/Llama-3.2-1B-Instruct (1.2B)
- Qwen/Qwen2.5-0.5B-Instruct (494M)
- openai/whisper-large-v3 (1.5B)
- Qwen/Qwen3-ASR-1.7B (2.3B)
- Qwen/Qwen3-1.7B (2.0B)
- EleutherAI/pythia-160m (213M)
- google/gemma-3-1b-it (1000M)
- baidu/Unlimited-OCR (3.3B)
- RadixArk/Kimi-K3-DSpark (2.2B)
- openai/whisper-small (242M)
- Qwen/Qwen3-VL-2B-Instruct (2.1B)
- Qwen/Qwen3.5-2B (2.3B)
- microsoft/Florence-2-base (232M)
- HuggingFaceTB/SmolLM2-135M (135M)
- MahmoudAshraf/mms-300m-1130-forced-aligner (315M)
- deepseek-ai/DeepSeek-OCR (3.3B)
- Qwen/Qwen3.5-0.8B (873M)
- facebook/sam3 (860M)
- zai-org/GLM-OCR (1.3B)
- Qwen/Qwen3-VL-Reranker-2B (2.1B)
- mlx-community/parakeet-tdt-0.6b-v2 (618M)
- prism-ml/Bonsai-27B-mlx-1bit (1.7B)
- Qwen/Qwen3-1.7B-Base (1.7B)
- Qwen/Qwen2-VL-2B-Instruct (2.2B)
- Qwen/Qwen2.5-0.5B (494M)
- mlx-community/parakeet-tdt-0.6b-v3 (627M)
- stabilityai/stable-diffusion-xl-base-1.0 (2.6B)
- vikhyatk/moondream2 (1.9B)
- gigant/romanian-wav2vec2 (315M)
- comodoro/wav2vec2-xls-r-300m-cs-250 (315M)
- nvidia/parakeet-ctc-1.1b (1.1B)
- stable-diffusion-v1-5/stable-diffusion-v1-5 (860M)
- HuggingFaceTB/SmolVLM2-500M-Video-Instruct (507M)
Models that fit at INT4 (quantized, 8 GB)
- google-bert/bert-base-uncased (110M)
- Qwen/Qwen3-0.6B (752M)
- openai-community/gpt2 (137M)
- Qwen/Qwen3-8B (8.2B)
- Qwen/Qwen3.5-9B (9.7B)
- Qwen/Qwen2.5-7B-Instruct (7.6B)
- Qwen/Qwen3-VL-8B-Instruct (8.8B)
- Qwen/Qwen2.5-VL-7B-Instruct (8.3B)
- Qwen/Qwen2.5-1.5B-Instruct (1.5B)
- Qwen/Qwen2.5-3B-Instruct (3.1B)
- Qwen/Qwen3.5-4B (4.7B)
- openai/whisper-large-v3-turbo (809M)
- Qwen/Qwen3-Embedding-0.6B (596M)
- meta-llama/Llama-3.2-1B-Instruct (1.2B)
- Qwen/Qwen3-4B (4.0B)
- Qwen/Qwen2.5-0.5B-Instruct (494M)
- meta-llama/Llama-3.1-8B-Instruct (8.0B)
- openai/whisper-large-v3 (1.5B)
- google/gemma-4-E4B-it (8.0B)
- Qwen/Qwen3-ASR-1.7B (2.3B)
- Qwen/Qwen3-VL-4B-Instruct (4.4B)
- Qwen/Qwen2.5-VL-3B-Instruct (3.8B)
- Qwen/Qwen3-1.7B (2.0B)
- EleutherAI/pythia-160m (213M)
- Qwen/Qwen3-4B-Instruct-2507 (4.0B)
- google/gemma-4-E2B-it (5.1B)
- google/gemma-4-12B-it (12.0B)
- Qwen/Qwen3-Embedding-4B (4.0B)
- google/gemma-3-1b-it (1000M)
- baidu/Unlimited-OCR (3.3B)
- RadixArk/Kimi-K3-DSpark (2.2B)
- openai/whisper-small (242M)
- datalab-to/chandra-ocr-2 (5.3B)
- Qwen/Qwen3-VL-2B-Instruct (2.1B)
- Qwen/Qwen3.5-2B (2.3B)
- microsoft/Florence-2-base (232M)
- HuggingFaceTB/SmolLM2-135M (135M)
- MahmoudAshraf/mms-300m-1130-forced-aligner (315M)
- deepseek-ai/DeepSeek-OCR (3.3B)
- Qwen/Qwen3.5-0.8B (873M)
Caveat: Requires a quantized checkpoint actually published for the model, check its Hugging Face page before relying on this.
All models that fit in 8 GB: the open models whose weights and overhead fit, at native, FP8 and INT4 precision.
Live pricing
No current live offer for this GPU, either there is no inventory listed right now, or the marketplace feed was briefly unreachable. This refreshes hourly; no price is shown rather than a stale or estimated one.