Best open translation models
How to choose an open machine translation model: language coverage, instruction following, model size, and license, with picks by situation.
Choose a translation model by the language pairs you need and how much control you want over the output. Coverage claims are easy to make and hard to check, so read exactly which languages a card lists, and run a few hundred of your own sentences through the shortlist before committing. Check that terminology, names and numbers survive translation, not only that the sentence reads well.
What matters for this task
Language coverage. The count of languages is less useful than whether your specific pair, in both directions, is covered. Low-resource pairs are where models diverge.
Instruction following. Newer models obey instructions about terminology, formatting and tone (for example keep product names, use formal register). This matters for business content more than raw fluency.
Off-target output. A common failure is answering in the wrong language. Benchmarks that report an off-target rate are worth reading.
Size and latency. Translation is high-volume, so small models that serve many requests per second often beat a large model. Mixture-of-experts models activate only part of their weights per token. See mixture of experts.
Quantization. Publishers sometimes ship quantized builds of the same model (the Hy-MT2 card lists FP8 and GGUF versions), which trade a little quality for lower memory. See quantization.
Quality metrics. COMET and similar scores compare against reference translations. Treat card numbers as the publisher's own evaluation.
Picks by situation
Small and fast. Hy-MT2-1.8B (Tencent, Apache 2.0, about 2.04B parameters in the catalog) belongs to a family covering 33 languages with translation-instruction following. The card says a 1.25-bit quantized build reduces the 1.8B model's storage to 440 MB for on-device use.
Balanced size. Hy-MT2-7B (Apache 2.0, about 8.03B) has the same 33-language coverage. The card says the 7B and 30B-A3B models outperform several larger open-source models on its evaluations.
Large with low active compute. Hy-MT2-30B-A3B (Apache 2.0, about 30.1B total parameters) is a larger translation model whose name marks it as an A3B variant. Check the card for its architecture and the memory it needs.
Widest language coverage. Index-Translate-9B (IndexTeam, Apache 2.0, about 9.65B parameters) is built on Qwen3.5 and supports text translation across 150 languages, following instructions on terminology, formatting, style and context. Its card reports 0.8789 on FLORES COMET-22 and publishes results for low-resource pairs, including off-target rates, as its own evaluation.
Licenses at a glance
All four picks above are Apache 2.0 per their model cards.
Running these on Aquanode
Rent a GPU by the hour and serve with a standard inference stack. See pricing and the VRAM calculator. A computed table of translation models renders below.
Open models for translation
All 6 models in the catalog for this task, grouped by size.
Under 3B parameters
| Model | Parameters | VRAM needed | License | Cheapest live fit | Est. $/hr |
|---|---|---|---|---|---|
| t5-3b | 2.9B | 12.7 GB | – | V100 | $0.088/hr |
| Hy-MT2-1.8B | 2.0B | 4.6 GB | – | RTX 4070 Super | $0.121/hr |
| Index-Translate-2B | 2.3B | 5.1 GB | – | RTX 4070 Super | $0.121/hr |
3B to 10B parameters
| Model | Parameters | VRAM needed | License | Cheapest live fit | Est. $/hr |
|---|---|---|---|---|---|
| Hy-MT2-7B | 8.0B | 17.9 GB | – | RTX A5000 | $0.176/hr |
| Index-Translate-9B | 9.7B | 21.6 GB | – | RTX A5000 | $0.176/hr |
10B to 40B parameters
| Model | Parameters | VRAM needed | License | Cheapest live fit | Est. $/hr |
|---|---|---|---|---|---|
| Hy-MT2-30B-A3B | 30.1B | 67.2 GB | – | A100 | $1.21/hr |
VRAM is for the precision each model is published in, with the same overhead and an 8,192-token context assumed on every page; see the methodology. The license column shows the license where our catalog records one. The fit is the lowest-priced single GPU type that holds the model at that precision, or the lowest-priced multi-GPU set (up to 8) when none does.
Sources
- https://huggingface.co/tencent/Hy-MT2-1.8B
- https://huggingface.co/tencent/Hy-MT2-7B
- https://huggingface.co/tencent/Hy-MT2-30B-A3B
- https://huggingface.co/IndexTeam/Index-Translate-9B
Updated 2026-10-07.
More ways to choose a model
- Best open models for coding
- Best open models for reasoning
- Best open models for chat and assistants
- Best open models for vision-language
- Best open models for OCR and document parsing
- Best open models for speech-to-text
- Best open models for text-to-speech
- Best open models for image generation
- Best open models for video generation
- Best open models for embeddings
By GPU memory: