Best open embedding models

How to choose an open text embedding model: embedding dimension, max tokens, multilingual coverage, Matryoshka sizes, and license, with picks by situation.

An embedding model turns text into a vector so you can search, cluster or rank by meaning. The two numbers that shape your system are the embedding dimension (how large your vector index is) and the maximum input tokens (how much of a document one vector can cover).

What matters for this task

Embedding dimension. Each vector has this many numbers. Bigger vectors can hold more information but cost more to store and search. Models trained with Matryoshka Representation Learning let you truncate vectors to smaller sizes with little loss.

Max tokens. Input beyond the limit is truncated. Short limits force you to chunk documents; long limits (tens of thousands of tokens) allow whole-document embeddings.

Multilingual and code coverage. Check the language list, and whether the model handles queries in one language against documents in another.

Instructions. Some models take a task instruction prefixed to the query, which the Qwen3 card says typically improves results by 1% to 5%.

Throughput. Embedding is a batch workload: a small model on one GPU can index millions of chunks. Model size mostly sets indexing time.

Quality. Compare on MTEB or your own retrieval set. Leaderboard positions on cards are the vendor's and are dated, so build a small set of real queries with known good answers and measure recall on it before you commit to a model and re-index a large corpus.

Picks by situation

Smallest with long context. voyage-4-nano (Voyage AI, Apache 2.0, about 346M parameters in the catalog) accepts a 32000 token context and supports 2048, 1024, 512 and 256 dimension embeddings, with 2048 the default.

Best balance, open license. Qwen3-Embedding-0.6B (Apache 2.0, about 0.6B) has a 32K token context and 1024 dimensions, adjustable from 32 to 1024. It covers 100+ languages including programming languages.

Highest quality. Qwen3-Embedding-8B (Apache 2.0, about 7.57B) gives 4096 dimensions (adjustable from 32 to 4096) with a 32K context. Its card says it ranked first on the MTEB multilingual leaderboard as of June 5, 2025. The middle option, Qwen3-Embedding-4B, uses 2560 dimensions and the same 32K context.

Short passages, MIT license. multilingual-e5-large-instruct (intfloat, MIT, about 0.56B) has 24 layers and 1024-dimension embeddings, supports 100 languages, and truncates input at 512 tokens, so plan for chunking.

Licenses at a glance

Apache 2.0: Qwen3-Embedding family, voyage-4-nano. MIT: multilingual-e5-large-instruct.

Running these on Aquanode

Rent a GPU by the hour to index a corpus, then keep serving queries on a smaller one. See pricing and the VRAM calculator. The table below lists all embedding models in the catalog.

Open models for embeddings

All 9 models in the catalog for this task, grouped by size.

Under 3B parameters

ModelParametersVRAM neededLicenseCheapest live fitEst. $/hr
Qwen3-Embedding-0.6B596M1.3 GBApache 2.0RTX 4070 Super$0.121/hr
multilingual-e5-large-instruct560M1.3 GBMITV100$0.088/hr
Qwen3-VL-Embedding-2B2.1B4.8 GB–RTX 4070 Super$0.121/hr
voyage-4-nano346M0.8 GB–RTX 4070 Super$0.121/hr
stella_en_1.5B_v51.5B6.9 GB–V100$0.088/hr

3B to 10B parameters

ModelParametersVRAM neededLicenseCheapest live fitEst. $/hr
Qwen3-Embedding-4B4.0B9.0 GB–RTX 4070 Super$0.121/hr
Qwen3-Embedding-8B7.6B16.9 GB–RTX A5000$0.176/hr
Qwen3-VL-Embedding-8B8.1B18.2 GB–RTX A5000$0.176/hr
pplx-embed-v2-context-9b-preview8.4B37.6 GB–RTX A6000$0.363/hr

VRAM is for the precision each model is published in, with the same overhead and an 8,192-token context assumed on every page; see the methodology. The license column shows the license where our catalog records one. The fit is the lowest-priced single GPU type that holds the model at that precision, or the lowest-priced multi-GPU set (up to 8) when none does.

Sources

Updated 2026-10-07.

More ways to choose a model

By GPU memory:

Submit the job. Everything after that is ours.

Sign up in 60 seconds. Pay for the GPU minutes you actually use.

© 2026 Aquanode. All rights reserved.

All trademarks, logos and brand names are the property of their respective owners.