Qwen / text-to-speech updated 6 months ago

Qwen3-TTS-12Hz-1.7B-VoiceDesign

&nbsp&nbsp🤗 Hugging Face&nbsp&nbsp &nbsp&nbsp🤖 ModelScope&nbsp&nbsp &nbsp&nbsp📑 Blog&nbsp&nbsp &nbsp&nbsp📑 Paper&nbsp&nbsp &nbsp&nbsp💻 GitHub

Params
1.9 B
Context
Downloads 30d
574 K
Likes
383
Commercial use: allowed apache-2.0 Not gated SAFETENSORS View on Hugging Face ↗

Download history

daily snapshots · 10 days
686 K574 K
Jul 28Jul 31Aug 3Aug 6

Can you run it?

Estimated VRAM at 8K context unless noted. Pick your hardware to see the verdict per quantization.

FileQuantSizeEst. VRAMVerdict on RTX 4090 · 24 GB
model.safetensors bf16 4.5 GB 5.8 GB ✅ Runs comfortably
Estimate: file size × 1.1 + KV cache at 8K + 0.5 GB overhead. Not a benchmark — how we calculate this.

Specifications

Architecture
Qwen3TTSForConditionalGeneration
Parameters
1.9 B
Tensor type
BF16
Licence
apache-2.0
First seen on the Hub
2026-01-21
Training datasets
undisclosed
Added to our catalog
2026-07-28