Best open speech-to-text models

How to choose an open speech recognition model: streaming latency, word error rate, language coverage, long audio, and license, with picks by situation.

Pick a speech-to-text model by deciding first whether you transcribe files after the fact or audio as it arrives. Everything else (size, languages, license) is secondary to that split, because streaming models are built differently from batch models.

What matters for this task

Streaming or offline. An offline model reads a whole clip, or a window of it, and returns text. A streaming model emits words while the speaker is still talking, which is what voice assistants and live captions need. Streaming models expose a delay setting: shorter delay means faster text and usually a higher word error rate.

Word error rate (WER). WER is the share of words the model gets wrong against a reference transcript. Compare it only on the same test set and language, and treat model-card numbers as the vendor's own measurement. Accents, noise and domain vocabulary move it more than parameter count does.

Languages. Coverage ranges from English-only checkpoints to dozens of languages. Check the exact list on the card, since a model that "supports" a language may be much weaker on it than on English.

Long audio and timestamps. Whisper has a 30-second receptive field, so longer files are processed as sliding or chunked windows (see the Whisper cards). Some models give word-level timestamps and punctuation out of the box, which matters for subtitles.

Memory. Most good ASR models are under 2B parameters, so they fit on modest GPUs; the cost driver is usually how many concurrent streams you serve. Use the VRAM calculator for your own numbers and see VRAM.

Picks by situation

Real-time voice agents. Voxtral Mini 4B Realtime 2602 (Mistral, Apache 2.0, about 4.4B parameters in the catalog) is described as a natively streaming model supporting 13 languages. Its card says it reaches accuracy comparable to offline systems with a delay under 500 ms, and that delay is configurable from 80 ms to 2.4 s, with 480 ms recommended.

Streaming across many languages. Nemotron 3.5 ASR Streaming 0.6B (NVIDIA, about 0.64B parameters) is a cache-aware streaming model covering 40 language-locales, with chunk sizes selectable at inference time from 80 ms to 1120 ms. Its license is OpenMDW 1.1, so read it before commercial use.

Fast batch transcription of European languages. Parakeet TDT 0.6B v3 (NVIDIA, CC-BY-4.0, 600M parameters) transcribes 25 European languages with automatic language detection, punctuation, capitalization and word-level timestamps. NVIDIA states it handles audio up to 24 minutes with full attention on an A100 80GB, or up to 3 hours with local attention.

General-purpose, widest ecosystem. Whisper large-v3 (OpenAI, Apache 2.0, about 1.54B parameters) is a multilingual encoder-decoder trained on a very large weakly supervised dataset. Whisper large-v3 turbo (MIT, about 0.81B) is the same model with decoder layers cut from 32 to 4; the card says this makes it much faster at the cost of a minor quality drop. Choose turbo when throughput matters more than the last bit of accuracy.

Chinese dialects and singing. Qwen3-ASR-1.7B (Qwen, Apache 2.0, about 2.35B parameters in the catalog) covers 30 languages plus 22 Chinese dialects, and its card lists streaming and offline inference from one model, including speech over background music. A companion forced-aligner model handles timestamps.

Licenses at a glance

Apache 2.0: Voxtral Realtime, Whisper large-v3, Qwen3-ASR. MIT: Whisper large-v3 turbo. CC-BY-4.0: Parakeet TDT v3. OpenMDW 1.1: Nemotron 3.5 ASR Streaming.

Running these on Aquanode

You can rent a GPU by the hour and serve any of these with its published inference code. See pricing for current rates. The table below lists every speech recognition model in the catalog with its size and license.

Open models for speech-to-text

All 51 models in the catalog for this task, grouped by size.

Under 3B parameters

ModelParametersVRAM neededLicenseCheapest live fitEst. $/hr
whisper-large-v3-turbo809M1.8 GB–V100$0.088/hr
whisper-large-v31.5B3.4 GBApache 2.0V100$0.088/hr
Qwen3-ASR-1.7B2.3B5.3 GB–RTX 4070 Super$0.121/hr
whisper-small242M1.1 GB–V100$0.088/hr
parakeet-ctc-1.1b1.1B4.8 GB–V100$0.088/hr
Qwen3-ASR-0.6B938M2.1 GB–RTX 4070 Super$0.121/hr
nemotron-3.5-asr-streaming-0.6b638M2.9 GB–V100$0.088/hr
parakeet-tdt-0.6b-v3627M2.8 GBCC BY 4.0V100$0.088/hr
distil-large-v3756M1.7 GB–V100$0.088/hr
Qwen3-ForcedAligner-0.6B918M2.1 GB–RTX 4070 Super$0.121/hr
cohere-transcribe-03-20262.1B4.6 GB–RTX 4070 Super$0.121/hr
whisper-medium764M3.4 GB–V100$0.088/hr
seamless-m4t-v2-large2.3B10.3 GB–V100$0.088/hr
granite-speech-4.1-2b2.3B5.2 GB–RTX 4070 Super$0.121/hr
mms-1b-all965M4.3 GB–V100$0.088/hr
nemotron-speech-streaming-en-0.6b618M2.8 GB–V100$0.088/hr
Qwen3-ASR-1.7B-hf2.0B4.6 GB–RTX 4070 Super$0.121/hr
diar_sortformer_4spk-v1124M0.6 GB–V100$0.088/hr
whisper-large-v21.5B6.9 GB–V100$0.088/hr
granite-speech-4.1-2b-plus2.1B4.7 GB–RTX 4070 Super$0.121/hr
medasr105M0.5 GB–V100$0.088/hr
Qwen3-ASR-0.6B-hf782M1.7 GB–RTX 4070 Super$0.121/hr
whisper-small.en242M1.1 GB–V100$0.088/hr
GLM-ASR-Nano-25122.3B5.0 GB–RTX 4070 Super$0.121/hr
parakeet-rnnt-0.6b617M2.8 GB–V100$0.088/hr
cohere-transcribe-arabic-07-20262.1B4.6 GB–RTX 4070 Super$0.121/hr
whisper-medium.en764M3.4 GB–V100$0.088/hr
whisper-large1.5B6.9 GB–V100$0.088/hr
granite-4.0-1b-speech2.3B5.2 GB–RTX 4070 Super$0.121/hr
CrisperWhisper2.0_large1.5B3.4 GB–RTX 4070 Super$0.121/hr
canary-qwen-2.5b2.6B5.7 GB–RTX 4070 Super$0.121/hr
Fun-ASR-Nano-2512-hf830M1.9 GB–RTX 4070 Super$0.121/hr
parakeet-ctc-0.6b609M2.7 GB–V100$0.088/hr
stt-2.6b-en-trfs2.7B6.0 GB–RTX 4070 Super$0.121/hr
BanglaASR242M1.1 GB–V100$0.088/hr
hviske-v5.32.1B4.6 GB–RTX 4070 Super$0.121/hr
VibeVoice-ASR-BitNet323M1.4 GB–V100$0.088/hr
nb-asr-beta-qwen06b-lunde05782M1.7 GB–RTX 4070 Super$0.121/hr
canary-1b-v2979M4.4 GB–V100$0.088/hr
distil-small.en166M0.4 GB–V100$0.088/hr
granite-speech-5.0-470m-turboctc473M1.1 GB–RTX 4070 Super$0.121/hr
granite-speech-5.0-470m-turboctc-nc473M1.1 GB–RTX 4070 Super$0.121/hr

3B to 10B parameters

ModelParametersVRAM neededLicenseCheapest live fitEst. $/hr
Voxtral-Mini-4B-Realtime-26024.4B9.9 GB–RTX 4070 Super$0.121/hr
VibeVoice-ASR8.7B19.4 GB–RTX A5000$0.176/hr
Phi-4-multimodal-instruct5.6B12.5 GB–RTX A4000$0.167/hr
granite-speech-3.3-2b3.0B6.7 GB–RTX 4070 Super$0.121/hr
granite-speech-3.2-8b8.5B19.0 GB–RTX A5000$0.176/hr
Audio8-ASR-Infinite4.1B9.1 GB–RTX 4070 Super$0.121/hr
granite-speech-3.3-8b8.6B19.3 GB–RTX A5000$0.176/hr
ARK-ASR-3B4.1B9.1 GB–RTX 4070 Super$0.121/hr
higgs-audio-v3-8b-stt-v28.9B19.9 GB–RTX A5000$0.176/hr

VRAM is for the precision each model is published in, with the same overhead and an 8,192-token context assumed on every page; see the methodology. The license column shows the license where our catalog records one. The fit is the lowest-priced single GPU type that holds the model at that precision, or the lowest-priced multi-GPU set (up to 8) when none does.

Sources

Updated 2026-10-07.

More ways to choose a model

By GPU memory:

Submit the job. Everything after that is ours.

Sign up in 60 seconds. Pay for the GPU minutes you actually use.

© 2026 Aquanode. All rights reserved.

All trademarks, logos and brand names are the property of their respective owners.