InternVL: OpenGVLab's open vision-language models
InternVL from OpenGVLab: InternVL2 and InternVL3 vision-language models, their LLM bases, licenses, and which to pick for document and image work.
What InternVL is
InternVL is a series of multimodal language models from OpenGVLab at Shanghai AI Lab. Each model follows a ViT-MLP-LLM layout: an InternViT vision encoder, an MLP projector, and a language model. The InternViT encoders come in 300M and 6B sizes, and the language models are borrowed from other families, including Qwen, InternLM, Phi-3 and Llama-3 variants according to the GitHub README. The project is released under the MIT license, with each model also subject to the license of the language model inside it.
Generations in order
The README timeline places InternVL 1.0 in December 2023 (InternViT-6B with a QLLaMA language component), InternVL 1.5 in April 2024 (4K image support and better OCR), InternVL 2.0 in July 2024, InternVL 2.5 in December 2024, and InternVL 3.0 in April 2025. The README also lists an InternVL 3.5 with a 241B-A28B top model and optimized 20B and 30B variants. We did not open a model card for that release, so we do not describe it further, and it has no hub of its own in this series.
InternVL2 (July 2024)
InternVL2 scaled the language component up to 34B. The InternVL2-8B card shows the pattern: an InternViT-300M-448px encoder, an MLP projector, and internlm2_5-7b-chat as the language model. It processes images at 448 pixel resolution, and the card states an 8k context window. It covers images, video and text, with document analysis, OCR, chart and infographic understanding. The card gives MIT for the project and Apache 2.0 for the InternLM component.
The README lists InternVL 2.5 (December 2024) with sizes from 1B to 78B. The hub holds the InternVL2 line and the InternVL2.5 checkpoints that sit with it, such as the 4B model.
InternVL3 (April 2025)
InternVL3 changed training more than layout. The InternVL3-8B card describes native multimodal pre-training, where vision and language are learned together instead of adding vision to a finished text model. Other documented changes:
- Language base: Qwen2.5-7B for the 8B model, which brings Apache 2.0 for that component alongside MIT for the project.
- Vision encoder: InternViT-300M-448px-V2_5, with 448x448 tiles at dynamic resolution and pixel unshuffle that cuts visual tokens to a quarter.
- Long context: Variable Visual Position Encoding (V2PE).
- Post-training: Mixed Preference Optimization for reasoning.
- Scope: multi-image and video input, GUI grounding, 3D vision and tool use.
The InternVL3 hub covers sizes from 1B up to 78B.
Which one to use today
- Starting a new project: InternVL3. It is the newer training recipe, with the same ViT-MLP-LLM layout, and it spans small to very large sizes.
- Small and cheap to serve: the 1B and 2B checkpoints in either hub. Smaller models trade accuracy for memory, and the model pages compute what each size needs on your hardware.
- Mid-size single-GPU work (OCR, charts, documents): InternVL3-8B or InternVL2-8B. InternVL3-8B is built on Qwen2.5-7B, InternVL2-8B on internlm2_5-7b-chat, so choose by which language base you prefer and which license you want to inherit.
- Largest open InternVL for hard multi-image reasoning: InternVL3-78B. The README says it achieves state-of-the-art performance in perception and reasoning, which is a vendor-reported claim we did not verify.
- Existing pipelines already tuned on InternVL2: staying on InternVL2 is reasonable, but check that your serving engine supports the version you move to before switching.
Whichever generation you choose, read the license of the language model inside the checkpoint as well as the MIT license on the project. The 8B cards we read pair MIT with Apache 2.0, but other sizes use different language bases, and those bases carry their own terms.
Side lines
InternVL has no separate side line in this series. The earlier InternVL 1.x checkpoints and community fine-tunes of InternVL2 and InternVL3 are indexed under the hubs, InternVL2 and InternVL3.
Running InternVL
You can run InternVL models on Aquanode GPUs.
InternVL generations
Every InternVL generation we track, oldest first. Each hub lists all of its models with VRAM at native, FP8 and INT4 precision.
| Hub | Models | Sizes | Smallest native VRAM | Cheapest live fit for it | Est. $/hr |
|---|---|---|---|---|---|
| InternVL2 | 4 | 938M to 25.5B | 2.1 GB | RTX 4070 Super | $0.121/hr |
| InternVL3 | 4 | 938M to 78.4B | 2.1 GB | RTX 4070 Super | $0.121/hr |
Smallest native VRAM is the lowest requirement among the publisher's own checkpoints in the hub, at the precision they are published in; see the methodology.
Best models by task
InternVL models appear on these ranked task pages.
Sources
- https://github.com/OpenGVLab/InternVL
- https://huggingface.co/OpenGVLab/InternVL3-8B
- https://huggingface.co/OpenGVLab/InternVL2-8B
Facts in the text above were read from these pages and are the publisher's own statements, not benchmarks run by Aquanode. Last reviewed 2026-10-07.