Best open text-to-speech models

How to choose an open text-to-speech model: voice cloning, real-time factor, languages, long-form audio, and license, with picks by situation.

Choose a text-to-speech model by what the voice has to do: speak one line in real time, clone a specific voice, or narrate a long multi-speaker script. Those are three different jobs, and few models are best at all of them.

What matters for this task

Real-time factor (RTF). RTF is generation time divided by audio length. Below 1.0 the model produces speech faster than it plays, which is the minimum for live use. Vendors quote RTF on specific hardware, so compare only like with like.

Streaming and time to first audio. For conversational agents, how soon the first sound plays matters more than total speed.

Voice cloning and voice design. Some models clone a voice from a short reference clip; others create a voice from a text description. Check the card for how much reference audio is needed and whether a transcript of it is required.

Languages. Coverage runs from English and Chinese only to dozens of languages. Unsupported languages can be unintelligible, so read the limits section of the card.

Length and speakers. Dialogue and podcast models handle several speakers and long scripts; single-line models do not.

License. Voice cloning raises misuse questions, and some licenses restrict commercial use or require consent for cloned voices.

Picks by situation

Open license, cloning, 30 languages. VoxCPM2 (OpenBMB, Apache 2.0, about 2.29B parameters) is a 2B model covering 30 languages with 48 kHz output, voice cloning, and voice design from a text description. The card reports an RTF around 0.3 on an RTX 4090 and about 0.13 with Nano-VLLM.

Streaming and the widest language list. Fish Audio S2 Pro (about 4.56B parameters in the catalog) is trained for 80+ languages with inline control of prosody and emotion. The card reports an RTF of 0.195 and about 100 ms to first audio on a single H200. Its license is the Fish Audio Research License, not an open-source license, so check terms before commercial use.

Two-speaker dialogue. Dia 1.6B (Nari Labs, Apache 2.0, about 1.61B parameters) generates dialogue directly from a transcript using [S1] and [S2] speaker tags, and can produce nonverbal sounds such as laughter. The card says it runs in real time on enterprise GPUs and slower on older ones.

Long multi-speaker audio. VibeVoice 1.5B (Microsoft, MIT, about 2.7B parameters in the catalog) synthesizes speech up to 90 minutes with up to 4 speakers. It is trained on English and Chinese only, and the card limits it to research use and prohibits unconsented voice impersonation.

Small and controllable by description. Parler-TTS Mini v1 (Apache 2.0, about 0.88B parameters) takes a text description of the voice and style, and ships 34 named speakers for consistency across generations. It is the lightest pick here.

Licenses at a glance

Apache 2.0: VoxCPM2, Dia, Parler-TTS Mini. MIT: VibeVoice 1.5B (with research-use limits in the card). Fish Audio Research License: S2 Pro.

Running these on Aquanode

Rent a GPU by the hour and run the vendor's inference code. See pricing for rates and the VRAM calculator for sizing. A computed table of every text-to-speech model in the catalog renders below.

Open models for text-to-speech

All 19 models in the catalog for this task, grouped by size.

Under 3B parameters

ModelParametersVRAM neededLicenseCheapest live fitEst. $/hr
VoxCPM22.3B5.1 GB–RTX 4070 Super$0.121/hr
VibeVoice-1.5B2.7B6.0 GB–RTX 4070 Super$0.121/hr
parler-tts-mini-multilingual-v1.1938M4.2 GB–V100$0.088/hr
stable-audio-3-small-sfx568M2.5 GB–V100$0.088/hr
Dia-1.6B1.6B7.2 GBApache 2.0V100$0.088/hr
MiniMax-Music32.4B10.9 GB–V100$0.088/hr
parler-tts-mini-v1878M3.9 GB–V100$0.088/hr
parler-tts-large-v12.3B10.4 GB–V100$0.088/hr
plapre-nano335M0.7 GB–RTX 4070 Super$0.121/hr
Audio8-TTS-Preview-0.1b170M0.4 GB–RTX 4070 Super$0.121/hr
sopro-v2-turbo122M0.5 GB–V100$0.088/hr
ice-012-audio714M1.6 GB–V100$0.088/hr

3B to 10B parameters

ModelParametersVRAM neededLicenseCheapest live fitEst. $/hr
s2-pro4.6B10.2 GB–RTX 4070 Super$0.121/hr
higgs-tts-3-4b4.7B10.4 GB–RTX 4070 Super$0.121/hr
YuE2-3B3.6B8.1 GB–RTX 4070 Super$0.121/hr
Veena3.8B8.5 GB–RTX 4070 Super$0.121/hr
acestep-5Hz-lm-4B4.2B9.4 GB–RTX 4070 Super$0.121/hr
Breeze-TTS-23.5B7.7 GB–RTX 4070 Super$0.121/hr
Irodori-TTS-v4-Large3.3B14.7 GB–V100$0.088/hr

VRAM is for the precision each model is published in, with the same overhead and an 8,192-token context assumed on every page; see the methodology. The license column shows the license where our catalog records one. The fit is the lowest-priced single GPU type that holds the model at that precision, or the lowest-priced multi-GPU set (up to 8) when none does.

Sources

Updated 2026-10-07.

More ways to choose a model

By GPU memory:

Submit the job. Everything after that is ours.

Sign up in 60 seconds. Pay for the GPU minutes you actually use.

© 2026 Aquanode. All rights reserved.

All trademarks, logos and brand names are the property of their respective owners.