1,011 models · refreshed nightly
Text to speech models
Every model in the catalog with its licence, estimated VRAM and daily-tracked downloads. Filters update the URL — share any view.
| # | Model | Params | Context | Commercial use | 30d | Min VRAM |
|---|---|---|---|---|---|---|
| 01 | Kokoro-82M | — | — | ✓ apache-2.0 | 12.0 M | — |
| 02 | XTTS-v2 | — | — | other | 9.2 M | — |
| 03 | chatterbox | — | — | ✓ mit | 2.5 M | from 12.2 GB |
| 04 | Qwen3-TTS-12Hz-1.7B-CustomVoice | 1.9 B | — | ✓ apache-2.0 | 2.3 M | from 5.8 GB |
| 05 | Qwen3-TTS-12Hz-0.6B-CustomVoice | 910 M | — | ✓ apache-2.0 | 1.6 M | from 3.4 GB |
| 06 | Kokoro-82M-v1.0-ONNX | — | — | ✓ apache-2.0 | 1.4 M | — |
| 07 | VoxCPM2 | 2.3 B | — | ✓ apache-2.0 | 900 K | from 5.9 GB |
| 08 | OmniVoice | 610 M | — | unknown | 877 K | from 4.2 GB |
| 09 | F5-TTS | — | — | ✗ cc-by-nc-4.0 | 770 K | from 5.0 GB |
| 10 | VibeVoice-Realtime-0.5B | 1.0 B | — | ✓ mit | 657 K | from 2.9 GB |
| 11 | Qwen3-TTS-12Hz-1.7B-VoiceDesign | 1.9 B | — | ✓ apache-2.0 | 574 K | from 5.8 GB |
| 12 | Qwen3-TTS-12Hz-0.6B-Base | 910 M | — | ✓ apache-2.0 | 445 K | from 3.4 GB |
| 13 | MOSS-TTS | 8.5 B | — | ✓ apache-2.0 | 375 K | from 20.5 GB |
| 14 | higgs-tts-3-4b | 4.7 B | — | other | 374 K | from 11.4 GB |
| 15 | VieNeu-TTS-v3-Turbo | 130 M | 1 K | ✓ apache-2.0 | 370 K | from 1.1 GB |
| 16 | s2-pro | 4.6 B | — | other | 368 K | from 11.2 GB |
VRAM figures are estimates for the smallest available quantization at 8K context — see /methodology. Downloads refresh nightly from the Hugging Face API.