Best open vision-language models

How to pick an open vision-language model: image input, context, license, and size. Picks by GPU budget with facts from the model cards.

A vision-language model takes images (and sometimes video) plus a text prompt and answers in text. Use one for chart and screenshot questions, document understanding, and visual agents. The catalog below lists every model tagged image-text-to-text; this page covers how to choose.

What matters when picking a vision-language model

Image tokens count against context. Every image is turned into tokens that sit in the context window next to your text. Many images or high-resolution pages use a lot of it, so context length matters more here than for text-only chat. The KV cache grows with it.

Native or bolted-on vision. Some models were trained on text and images together, others add a vision encoder to an existing language model. Check what the card says about training and which tasks it reports.

Thinking mode. Some current models think before answering by default, which improves hard visual reasoning and adds output tokens. Check whether you can switch it off for latency.

Size and architecture. A mixture-of-experts model stores all of its weights but runs a few per token, which makes it fast for its quality. You still need memory for the full weights.

License. Check the card before commercial use.

Precision. BF16 is the reference; a quantized build lowers the footprint. Size it in the VRAM calculator.

Picks by situation

Smaller GPUs (a 24 GB card): Qwen3.5-9B. The card lists 9B parameters, a 262,144-token native context extensible to 1,010,000, and Apache 2.0. Qwen says it uses early-fusion training on multimodal tokens, and that Qwen3.5 operates in thinking mode by default.

Efficient mixture-of-experts: Kimi-VL-A3B-Instruct. Moonshot lists 16B total and 3B activated parameters, a 128K context, and an MIT license. The card says it handles OCR, multi-image input, and video, and uses a native-resolution vision encoder (MoonViT). A thinking variant exists as a separate model.

One 80 GB GPU: Qwen3.5-27B. 27B parameters, the same 262,144-token native context, and Apache 2.0 per the card. Qwen3.5-35B-A3B is the mixture-of-experts option in the same family: 35B in total and 3B activated, with the same context and license. Qwen reports in the card that Qwen3.5 outperforms Qwen3-VL across reasoning, coding, agents, and visual understanding benchmarks (vendor-reported).

Longest context: the Qwen3.5 models above, at 262,144 tokens natively and up to 1,010,000 with extension, per their cards.

Most permissive license: Qwen3.5 is Apache 2.0 and Kimi-VL-A3B-Instruct is MIT, per their cards.

Serving notes

Resize or tile images before sending them: more pixels mean more tokens, more memory, and slower responses. Batch similar request shapes together. For document pages specifically, a dedicated model from the OCR page is often cheaper than a general vision-language model.

Aquanode rents GPUs by the hour, so you can test a 9B model and a larger one on the right cards and keep whichever you need. See pricing. Benchmark claims on cards are vendor-reported; test on your own images.

Open models for vision-language

The 60 most downloaded of 144 models in the catalog, grouped by size. The rest are listed on the models directory.

Under 3B parameters

ModelParametersVRAM neededLicenseCheapest live fitEst. $/hr
Qwen3-VL-2B-Instruct2.1B4.8 GB–RTX 4070 Super$0.121/hr
Qwen3.5-2B2.3B5.1 GB–RTX 4070 Super$0.121/hr
Florence-2-base232M0.5 GB–V100$0.088/hr
Qwen3.5-0.8B873M2.0 GB–RTX 4070 Super$0.121/hr
GLM-OCR1.3B3.0 GB–RTX 4070 Super$0.121/hr
Qwen2-VL-2B-Instruct2.2B4.9 GB–RTX 4070 Super$0.121/hr
moondream21.9B4.3 GB–RTX 4070 Super$0.121/hr
SmolVLM2-500M-Video-Instruct507M2.3 GB–V100$0.088/hr
Qwen3.5-2B-Base2.3B5.1 GB–RTX 4070 Super$0.121/hr
surya-ocr-2686M1.5 GB–RTX 4070 Super$0.121/hr
Cosmos-Reason2-2B2.4B5.5 GB–RTX 4070 Super$0.121/hr
Rax-4.52.3B5.1 GB–RTX 4070 Super$0.121/hr
GOT-OCR2_0716M1.6 GB–RTX 4070 Super$0.121/hr
HunyuanOCR1.1B2.5 GB–RTX 4070 Super$0.121/hr
Florence-2-large777M1.7 GB–V100$0.088/hr
llava-onevision-qwen2-0.5b-ov-hf894M2.0 GB–V100$0.088/hr
InternVL2-2B2.2B4.9 GB–RTX 4070 Super$0.121/hr
InternVL2-1B938M2.1 GB–RTX 4070 Super$0.121/hr
SmolVLM-256M-Instruct256M0.6 GB–RTX 4070 Super$0.121/hr
MiniCPM-V-4.61.3B2.9 GB–RTX 4070 Super$0.121/hr

3B to 10B parameters

ModelParametersVRAM neededLicenseCheapest live fitEst. $/hr
Qwen3.5-9B9.7B21.6 GBApache 2.0RTX A5000$0.176/hr
Qwen3-VL-8B-Instruct8.8B19.6 GB–RTX A5000$0.176/hr
Qwen2.5-VL-7B-Instruct8.3B18.5 GB–RTX A5000$0.176/hr
Qwen3.5-4B4.7B10.4 GB–RTX 4070 Super$0.121/hr
Qwen3-VL-4B-Instruct4.4B9.9 GB–RTX 4070 Super$0.121/hr
Qwen2.5-VL-3B-Instruct3.8B8.4 GB–RTX 4070 Super$0.121/hr
Unlimited-OCR3.3B7.5 GB–RTX 4070 Super$0.121/hr
chandra-ocr-25.3B11.8 GB–RTX 4070 Super$0.121/hr
DeepSeek-OCR3.3B7.5 GBMITRTX 4070 Super$0.121/hr
llava-1.5-7b-hf7.1B15.8 GB–V100$0.088/hr
gemma-3-4b-it4.3B9.6 GBGemma Terms of UseRTX 4070 Super$0.121/hr
Qwen2-VL-7B-Instruct8.3B18.5 GB–RTX A5000$0.176/hr
DeepSeek-OCR-23.4B7.6 GB–RTX 4070 Super$0.121/hr
medgemma-4b-it4.3B9.6 GB–RTX 4070 Super$0.121/hr
Phi-3.5-vision-instruct4.1B9.3 GB–RTX 4070 Super$0.121/hr
UI-TARS-1.5-7B8.3B37.1 GB–RTX A6000$0.363/hr
llava-v1.6-mistral-7b-hf7.6B16.9 GB–RTX A5000$0.176/hr
Cosmos-Reason2-8B8.8B19.6 GB–RTX A5000$0.176/hr
dots.mocr3.0B6.8 GB–RTX 4070 Super$0.121/hr
blip2-opt-2.7b3.7B16.7 GB–RTX A5000$0.176/hr
Qwen3.5-4B-Base4.7B10.4 GB–RTX 4070 Super$0.121/hr
Qwen3.5-9B-Base9.7B21.6 GB–RTX A5000$0.176/hr

10B to 40B parameters

ModelParametersVRAM neededLicenseCheapest live fitEst. $/hr
gemma-4-31B-it31.3B69.9 GBApache 2.0A100$1.21/hr
gemma-4-26B-A4B-it25.8B57.7 GBApache 2.0A100$1.21/hr
Qwen3.6-27B27.8B62.1 GB–A100$1.21/hr
Qwen3.8-27B27.8B62.1 GB–A100$1.21/hr
Qwen3.6-35B-A3B36.0B80.4 GBApache 2.0RTX PRO 6000$1.38/hr
Qwen3.5-27B27.8B62.1 GB–A100$1.21/hr
Qwen3.5-35B-A3B36.0B80.4 GB–RTX PRO 6000$1.38/hr
Qwen2.5-VL-32B-Instruct33.5B74.8 GB–A100$1.21/hr
diffusiongemma-26B-A4B-it25.8B57.7 GB–A100$1.21/hr
gemma-3-12b-it12.2B27.2 GBGemma Terms of UseRTX A6000$0.363/hr
gemma-4-31B32.7B73.1 GB–A100$1.21/hr
Muse-Glimmer-30B29.8B66.6 GB–A100$1.21/hr
Qwen3-VL-32B-Instruct33.4B74.6 GBApache 2.0A100$1.21/hr
gemma-3-27b-it27.4B61.3 GB–A100$1.21/hr
Infinity-Parser2-Pro35.1B78.5 GB–A100$1.21/hr

40B to 150B parameters

ModelParametersVRAM neededLicenseCheapest live fitEst. $/hr
Qwen3.5-122B-A10B125.1B280 GBApache 2.0RTX A6000 × 6$2.18/hr

150B parameters and up

ModelParametersVRAM neededLicenseCheapest live fitEst. $/hr
Qwen3-VL-235B-A22B-Instruct235.7B527 GB–RTX PRO 6000 × 6$8.25/hr
GLM-5.3-Flash321.3B359 GB–RTX PRO 6000 × 4$5.50/hr

VRAM is for the precision each model is published in, with the same overhead and an 8,192-token context assumed on every page; see the methodology. The license column shows the license where our catalog records one. The fit is the lowest-priced single GPU type that holds the model at that precision, or the lowest-priced multi-GPU set (up to 8) when none does. See all models for the rest.

Sources

Updated 2026-10-07.

More ways to choose a model

By GPU memory:

Submit the job. Everything after that is ours.

Sign up in 60 seconds. Pay for the GPU minutes you actually use.

© 2026 Aquanode. All rights reserved.

All trademarks, logos and brand names are the property of their respective owners.