1,054 models · refreshed nightly

Text to speech models

Every model in the catalog with its licence, estimated VRAM and daily-tracked downloads. Filters update the URL — share any view.

#ModelParamsContextCommercial use30dMin VRAM
01 Kokoro-82M hexgrad · text-to-speech 82 M ✓ apache-2.0 11.6 M from 0.7 GB
02 XTTS-v2 coqui · text-to-speech other 7.3 M
03 audio.cpp-gguf audio-cpp · text-to-speech other 4.3 M from 1.4 GB
04 Qwen3-TTS-12Hz-1.7B-CustomVoice Qwen · text-to-speech 1.9 B ✓ apache-2.0 2.6 M from 5.8 GB
05 chatterbox ResembleAI · text-to-speech ✓ mit 1.8 M from 12.2 GB
06 OmniVoice k2-fsa · text-to-speech 610 M unknown 1.3 M from 4.2 GB
07 Qwen3-TTS-12Hz-0.6B-CustomVoice Qwen · text-to-speech 910 M ✓ apache-2.0 1.0 M from 3.4 GB
08 F5-TTS SWivid · text-to-speech ✗ cc-by-nc-4.0 949 K from 5.0 GB
09 Kokoro-82M-v1.0-ONNX onnx-community · text-to-speech 82 M ✓ apache-2.0 726 K from 0.7 GB
10 Qwen3-TTS-12Hz-0.6B-Base Qwen · text-to-speech 910 M ✓ apache-2.0 648 K from 3.4 GB
11 VieNeu-TTS-v3-Turbo pnnbao-ump · text-to-speech 130 M 1 K ✓ apache-2.0 601 K from 1.1 GB
12 VibeVoice-1.5B microsoft · text-to-speech 2.7 B ✓ mit 482 K from 6.9 GB
VRAM figures are estimates for the smallest available quantization at 8K context — see /methodology. Downloads refresh nightly from the Hugging Face API.