Qwen3-TTS-12Hz-0.6B-Base
Qwen3-TTS is a family of advanced multilingual, controllable, robust, and streaming text-to-speech models. Trained on over 5 million hours of speech data spanning 10 languages, Qwen3-TTS supports state-of-the-art 3-second voice cloning and description-based control.
Params
910 M
Context
—
Downloads 30d
648 K
Likes
297
Download history
daily snapshots · 55 days
▲ 113 K in the last 30 days (21.1%)
648 K483 K
Aug 22Sep 1Sep 11Sep 20
648 K410 K
Jul 28Aug 15Sep 2Sep 20
Can you run it?
Estimated VRAM at 8K context unless noted. Pick your hardware to see the verdict per quantization.
| File | Quant | Size | Est. VRAM | Verdict on RTX 4090 · 24 GB |
|---|---|---|---|---|
| model.safetensors | bf16 | 2.5 GB | 3.4 GB | ✅ Runs comfortably |
Estimate: file size × 1.1 + KV cache at 8K + 0.5 GB overhead. Not a benchmark — how we calculate this.
Specifications
- Architecture
- Qwen3TTSForConditionalGeneration
- Parameters
- 910 M
- Tensor type
- BF16
- Licence
- apache-2.0
- First seen on the Hub
- 2026-01-21
- Training datasets
- undisclosed
- Added to our catalog
- 2026-07-28
Compare with any text-to-speech model
Popular comparisons
Appears in