Best open text-to-speech models
How to choose an open text-to-speech model: voice cloning, real-time factor, languages, long-form audio, and license, with picks by situation.
Choose a text-to-speech model by what the voice has to do: speak one line in real time, clone a specific voice, or narrate a long multi-speaker script. Those are three different jobs, and few models are best at all of them.
What matters for this task
Real-time factor (RTF). RTF is generation time divided by audio length. Below 1.0 the model produces speech faster than it plays, which is the minimum for live use. Vendors quote RTF on specific hardware, so compare only like with like.
Streaming and time to first audio. For conversational agents, how soon the first sound plays matters more than total speed.
Voice cloning and voice design. Some models clone a voice from a short reference clip; others create a voice from a text description. Check the card for how much reference audio is needed and whether a transcript of it is required.
Languages. Coverage runs from English and Chinese only to dozens of languages. Unsupported languages can be unintelligible, so read the limits section of the card.
Length and speakers. Dialogue and podcast models handle several speakers and long scripts; single-line models do not.
License. Voice cloning raises misuse questions, and some licenses restrict commercial use or require consent for cloned voices.
Picks by situation
Open license, cloning, 30 languages. VoxCPM2 (OpenBMB, Apache 2.0, about 2.29B parameters) is a 2B model covering 30 languages with 48 kHz output, voice cloning, and voice design from a text description. The card reports an RTF around 0.3 on an RTX 4090 and about 0.13 with Nano-VLLM.
Streaming and the widest language list. Fish Audio S2 Pro (about 4.56B parameters in the catalog) is trained for 80+ languages with inline control of prosody and emotion. The card reports an RTF of 0.195 and about 100 ms to first audio on a single H200. Its license is the Fish Audio Research License, not an open-source license, so check terms before commercial use.
Two-speaker dialogue. Dia 1.6B (Nari Labs, Apache 2.0, about 1.61B parameters) generates dialogue directly from a transcript using [S1] and [S2] speaker tags, and can produce nonverbal sounds such as laughter. The card says it runs in real time on enterprise GPUs and slower on older ones.
Long multi-speaker audio. VibeVoice 1.5B (Microsoft, MIT, about 2.7B parameters in the catalog) synthesizes speech up to 90 minutes with up to 4 speakers. It is trained on English and Chinese only, and the card limits it to research use and prohibits unconsented voice impersonation.
Small and controllable by description. Parler-TTS Mini v1 (Apache 2.0, about 0.88B parameters) takes a text description of the voice and style, and ships 34 named speakers for consistency across generations. It is the lightest pick here.
Licenses at a glance
Apache 2.0: VoxCPM2, Dia, Parler-TTS Mini. MIT: VibeVoice 1.5B (with research-use limits in the card). Fish Audio Research License: S2 Pro.
Running these on Aquanode
Rent a GPU by the hour and run the vendor's inference code. See pricing for rates and the VRAM calculator for sizing. A computed table of every text-to-speech model in the catalog renders below.
Open models for text-to-speech
All 19 models in the catalog for this task, grouped by size.
Under 3B parameters
| Model | Parameters | VRAM needed | License | Cheapest live fit | Est. $/hr |
|---|---|---|---|---|---|
| VoxCPM2 | 2.3B | 5.1 GB | – | RTX 4070 Super | $0.121/hr |
| VibeVoice-1.5B | 2.7B | 6.0 GB | – | RTX 4070 Super | $0.121/hr |
| parler-tts-mini-multilingual-v1.1 | 938M | 4.2 GB | – | V100 | $0.088/hr |
| stable-audio-3-small-sfx | 568M | 2.5 GB | – | V100 | $0.088/hr |
| Dia-1.6B | 1.6B | 7.2 GB | Apache 2.0 | V100 | $0.088/hr |
| MiniMax-Music3 | 2.4B | 10.9 GB | – | V100 | $0.088/hr |
| parler-tts-mini-v1 | 878M | 3.9 GB | – | V100 | $0.088/hr |
| parler-tts-large-v1 | 2.3B | 10.4 GB | – | V100 | $0.088/hr |
| plapre-nano | 335M | 0.7 GB | – | RTX 4070 Super | $0.121/hr |
| Audio8-TTS-Preview-0.1b | 170M | 0.4 GB | – | RTX 4070 Super | $0.121/hr |
| sopro-v2-turbo | 122M | 0.5 GB | – | V100 | $0.088/hr |
| ice-012-audio | 714M | 1.6 GB | – | V100 | $0.088/hr |
3B to 10B parameters
| Model | Parameters | VRAM needed | License | Cheapest live fit | Est. $/hr |
|---|---|---|---|---|---|
| s2-pro | 4.6B | 10.2 GB | – | RTX 4070 Super | $0.121/hr |
| higgs-tts-3-4b | 4.7B | 10.4 GB | – | RTX 4070 Super | $0.121/hr |
| YuE2-3B | 3.6B | 8.1 GB | – | RTX 4070 Super | $0.121/hr |
| Veena | 3.8B | 8.5 GB | – | RTX 4070 Super | $0.121/hr |
| acestep-5Hz-lm-4B | 4.2B | 9.4 GB | – | RTX 4070 Super | $0.121/hr |
| Breeze-TTS-2 | 3.5B | 7.7 GB | – | RTX 4070 Super | $0.121/hr |
| Irodori-TTS-v4-Large | 3.3B | 14.7 GB | – | V100 | $0.088/hr |
VRAM is for the precision each model is published in, with the same overhead and an 8,192-token context assumed on every page; see the methodology. The license column shows the license where our catalog records one. The fit is the lowest-priced single GPU type that holds the model at that precision, or the lowest-priced multi-GPU set (up to 8) when none does.
Sources
- https://huggingface.co/openbmb/VoxCPM2
- https://huggingface.co/fishaudio/s2-pro
- https://huggingface.co/nari-labs/Dia-1.6B
- https://huggingface.co/parler-tts/parler-tts-mini-v1
- https://huggingface.co/microsoft/VibeVoice-1.5B
Updated 2026-10-07.
More ways to choose a model
- Best open models for coding
- Best open models for reasoning
- Best open models for chat and assistants
- Best open models for vision-language
- Best open models for OCR and document parsing
- Best open models for speech-to-text
- Best open models for image generation
- Best open models for video generation
- Best open models for embeddings
- Best open models for translation
By GPU memory: