Qwen3-TTS-12Hz-1.7B-CustomVoice
Qwen3-TTS covers 10 major languages (Chinese, English, Japanese, Korean, German, French, Russian, Portuguese, Spanish, and Italian) as well as multiple dialectal voice profiles to meet global application needs. In addition, the models feature strong contextual understanding, enabling adaptive control of tone, speaking rate, and emotional expression based on instructions and text semantics, and they show markedly impr...
Params
1.9 B
Context
—
Downloads 30d
2.6 M
Likes
1,979
Download history
daily snapshots · 55 days
▲ 320 K in the last 30 days (14.2%)
2.6 M2.3 M
Aug 22Sep 1Sep 11Sep 20
2.6 M2.1 M
Jul 28Aug 15Sep 2Sep 20
Can you run it?
Estimated VRAM at 8K context unless noted. Pick your hardware to see the verdict per quantization.
| File | Quant | Size | Est. VRAM | Verdict on RTX 4090 · 24 GB |
|---|---|---|---|---|
| model.safetensors | bf16 | 4.5 GB | 5.8 GB | ✅ Runs comfortably |
Estimate: file size × 1.1 + KV cache at 8K + 0.5 GB overhead. Not a benchmark — how we calculate this.
Specifications
- Architecture
- Qwen3TTSForConditionalGeneration
- Parameters
- 1.9 B
- Tensor type
- BF16
- Licence
- apache-2.0
- First seen on the Hub
- 2026-01-21
- Training datasets
- undisclosed
- Added to our catalog
- 2026-07-28
Compare with any text-to-speech model
Popular comparisons
Appears in