1,011 models · refreshed nightly

Text to speech models

Every model in the catalog with its licence, estimated VRAM and daily-tracked downloads. Filters update the URL — share any view.

#ModelParamsContextCommercial use30dMin VRAM
01 Kokoro-82M hexgrad · text-to-speech ✓ apache-2.0 12.0 M
02 XTTS-v2 coqui · text-to-speech other 9.2 M
03 chatterbox ResembleAI · text-to-speech ✓ mit 2.5 M from 12.2 GB
04 Qwen3-TTS-12Hz-1.7B-CustomVoice Qwen · text-to-speech 1.9 B ✓ apache-2.0 2.3 M from 5.8 GB
05 Qwen3-TTS-12Hz-0.6B-CustomVoice Qwen · text-to-speech 910 M ✓ apache-2.0 1.6 M from 3.4 GB
06 Kokoro-82M-v1.0-ONNX onnx-community · text-to-speech ✓ apache-2.0 1.4 M
07 VoxCPM2 openbmb · text-to-speech 2.3 B ✓ apache-2.0 900 K from 5.9 GB
08 OmniVoice k2-fsa · text-to-speech 610 M unknown 877 K from 4.2 GB
09 F5-TTS SWivid · text-to-speech ✗ cc-by-nc-4.0 770 K from 5.0 GB
10 VibeVoice-Realtime-0.5B microsoft · text-to-speech 1.0 B ✓ mit 657 K from 2.9 GB
11 Qwen3-TTS-12Hz-1.7B-VoiceDesign Qwen · text-to-speech 1.9 B ✓ apache-2.0 574 K from 5.8 GB
12 Qwen3-TTS-12Hz-0.6B-Base Qwen · text-to-speech 910 M ✓ apache-2.0 445 K from 3.4 GB
13 MOSS-TTS OpenMOSS-Team · text-to-speech 8.5 B ✓ apache-2.0 375 K from 20.5 GB
14 higgs-tts-3-4b bosonai · text-to-speech 4.7 B other 374 K from 11.4 GB
15 VieNeu-TTS-v3-Turbo pnnbao-ump · text-to-speech 130 M 1 K ✓ apache-2.0 370 K from 1.1 GB
16 s2-pro fishaudio · text-to-speech 4.6 B other 368 K from 11.2 GB
VRAM figures are estimates for the smallest available quantization at 8K context — see /methodology. Downloads refresh nightly from the Hugging Face API.