InternVL: OpenGVLab's open vision-language models

InternVL from OpenGVLab: InternVL2 and InternVL3 vision-language models, their LLM bases, licenses, and which to pick for document and image work.

What InternVL is

InternVL is a series of multimodal language models from OpenGVLab at Shanghai AI Lab. Each model follows a ViT-MLP-LLM layout: an InternViT vision encoder, an MLP projector, and a language model. The InternViT encoders come in 300M and 6B sizes, and the language models are borrowed from other families, including Qwen, InternLM, Phi-3 and Llama-3 variants according to the GitHub README. The project is released under the MIT license, with each model also subject to the license of the language model inside it.

Generations in order

The README timeline places InternVL 1.0 in December 2023 (InternViT-6B with a QLLaMA language component), InternVL 1.5 in April 2024 (4K image support and better OCR), InternVL 2.0 in July 2024, InternVL 2.5 in December 2024, and InternVL 3.0 in April 2025. The README also lists an InternVL 3.5 with a 241B-A28B top model and optimized 20B and 30B variants. We did not open a model card for that release, so we do not describe it further, and it has no hub of its own in this series.

InternVL2 (July 2024)

InternVL2 scaled the language component up to 34B. The InternVL2-8B card shows the pattern: an InternViT-300M-448px encoder, an MLP projector, and internlm2_5-7b-chat as the language model. It processes images at 448 pixel resolution, and the card states an 8k context window. It covers images, video and text, with document analysis, OCR, chart and infographic understanding. The card gives MIT for the project and Apache 2.0 for the InternLM component.

The README lists InternVL 2.5 (December 2024) with sizes from 1B to 78B. The hub holds the InternVL2 line and the InternVL2.5 checkpoints that sit with it, such as the 4B model.

InternVL3 (April 2025)

InternVL3 changed training more than layout. The InternVL3-8B card describes native multimodal pre-training, where vision and language are learned together instead of adding vision to a finished text model. Other documented changes:

  • Language base: Qwen2.5-7B for the 8B model, which brings Apache 2.0 for that component alongside MIT for the project.
  • Vision encoder: InternViT-300M-448px-V2_5, with 448x448 tiles at dynamic resolution and pixel unshuffle that cuts visual tokens to a quarter.
  • Long context: Variable Visual Position Encoding (V2PE).
  • Post-training: Mixed Preference Optimization for reasoning.
  • Scope: multi-image and video input, GUI grounding, 3D vision and tool use.

The InternVL3 hub covers sizes from 1B up to 78B.

Which one to use today

  • Starting a new project: InternVL3. It is the newer training recipe, with the same ViT-MLP-LLM layout, and it spans small to very large sizes.
  • Small and cheap to serve: the 1B and 2B checkpoints in either hub. Smaller models trade accuracy for memory, and the model pages compute what each size needs on your hardware.
  • Mid-size single-GPU work (OCR, charts, documents): InternVL3-8B or InternVL2-8B. InternVL3-8B is built on Qwen2.5-7B, InternVL2-8B on internlm2_5-7b-chat, so choose by which language base you prefer and which license you want to inherit.
  • Largest open InternVL for hard multi-image reasoning: InternVL3-78B. The README says it achieves state-of-the-art performance in perception and reasoning, which is a vendor-reported claim we did not verify.
  • Existing pipelines already tuned on InternVL2: staying on InternVL2 is reasonable, but check that your serving engine supports the version you move to before switching.

Whichever generation you choose, read the license of the language model inside the checkpoint as well as the MIT license on the project. The 8B cards we read pair MIT with Apache 2.0, but other sizes use different language bases, and those bases carry their own terms.

Side lines

InternVL has no separate side line in this series. The earlier InternVL 1.x checkpoints and community fine-tunes of InternVL2 and InternVL3 are indexed under the hubs, InternVL2 and InternVL3.

Running InternVL

You can run InternVL models on Aquanode GPUs.

InternVL generations

Every InternVL generation we track, oldest first. Each hub lists all of its models with VRAM at native, FP8 and INT4 precision.

HubModelsSizesSmallest native VRAMCheapest live fit for itEst. $/hr
InternVL24938M to 25.5B2.1 GBRTX 4070 Super$0.121/hr
InternVL34938M to 78.4B2.1 GBRTX 4070 Super$0.121/hr

Smallest native VRAM is the lowest requirement among the publisher's own checkpoints in the hub, at the precision they are published in; see the methodology.

Best models by task

InternVL models appear on these ranked task pages.

Sources

Facts in the text above were read from these pages and are the publisher's own statements, not benchmarks run by Aquanode. Last reviewed 2026-10-07.

Submit the job. Everything after that is ours.

Sign up in 60 seconds. Pay for the GPU minutes you actually use.

© 2026 Aquanode. All rights reserved.

All trademarks, logos and brand names are the property of their respective owners.