Audio Text To Text Models

Tracked audio text to text models monitored by OpenModelStats. Rankings below compare tracked models against each other — see the methodology for scope and limitations.

Top-10 combined DL 30D
2M
Tasks directory position
See all tasks
Growth history
Accumulating
Last refresh
Aug 24, 2026

Trending now

#ModelDL 30D7D
1MOSS-Transcribe-DiarizeOpenMOSS-Team281.4K
2audio-flamingo-3nvidia265
3audio-flamingo-3-hfnvidia135.1K
4ultravox-v0_7-glm-4_6fixie-ai1,771
5MOSS-Music-8B-InstructOpenMOSS-Team1,566
6music-flamingo-2601-hfnvidia25.5K
7NemotronLabs-VoiceChat-11B-MXFP8OsaurusAI79
8shuka-1sarvamai305
9Voxtral-Small-24B-2507mistralai253.1K
10ultravox-v0_5-llama-3_2-1bfixie-ai593.1K

Most downloaded

New & notable

Published within the last 90 days with meaningful traction.

ModelDL 30D7D gained
Qwen2-Audio_PCLM_DPOIHP-Lab3,368
MOSS-Audio-4B-Instruct-GGUFcstr1,525
mistralai_Voxtral-Small-24B-2507-GGUFAbu-Dju1,467
Audio Text To Text Models — Most Downloaded & Trending | OpenModelStats