Qwen / automatic-speech-recognition updated 6 months ago

Qwen3-ForcedAligner-0.6B

The Qwen3-ASR family includes Qwen3-ASR-1.7B and Qwen3-ASR-0.6B, which support language identification and ASR for 52 languages and dialects. Both leverage large-scale speech training data and the strong audio understanding capability of their foundation model, Qwen3-Omni. Experiments show that the 1.7B version achieves state-of-the-art performance among open-source ASR models and is competitive with the strongest pr...

Params
920 M
Context
Downloads 30d
407 K
Likes
150
Commercial use: allowed apache-2.0 Not gated SAFETENSORS View on Hugging Face ↗

Download history

daily snapshots · 10 days
421 K399 K
Jul 28Jul 31Aug 3Aug 6

Can you run it?

Estimated VRAM at 8K context unless noted. Pick your hardware to see the verdict per quantization.

FileQuantSizeEst. VRAMVerdict on RTX 4090 · 24 GB
model.safetensors bf16 1.8 GB 2.7 GB ✅ Runs comfortably
Estimate: file size × 1.1 + KV cache at 8K + 0.5 GB overhead. Not a benchmark — how we calculate this.

Specifications

Architecture
Qwen3ASRForConditionalGeneration
Parameters
920 M
Tensor type
BF16
Licence
apache-2.0
First seen on the Hub
2026-01-28
Training datasets
undisclosed
Added to our catalog
2026-07-28