microsoft / automatic-speech-recognition updated 6 months ago

VibeVoice-ASR

VibeVoice-ASR is a unified speech-to-text model designed to handle 60-minute long-form audio in a single pass, generating structured transcriptions containing Who (Speaker), When (Timestamps), and What (Content), with support for Customized Hotwords and over 50 languages.

Params
8.7 B
Context
Downloads 30d
695 K
Likes
1,253
Commercial use: allowed mit Not gated SAFETENSORS 51 languages View on Hugging Face ↗

Download history

daily snapshots · 10 days
696 K679 K
Jul 28Jul 31Aug 3Aug 6

Can you run it?

Estimated VRAM at 8K context unless noted. Pick your hardware to see the verdict per quantization.

FileQuantSizeEst. VRAMVerdict on RTX 4090 · 24 GB
model.safetensors bf16 17.3 GB 20.9 GB ⚠️ Tight — reduce context
Estimate: file size × 1.1 + KV cache at 8K + 0.5 GB overhead. Not a benchmark — how we calculate this.

Specifications

Architecture
VibeVoiceForASRTraining
Parameters
8.7 B
Tensor type
BF16
Licence
mit
First seen on the Hub
2026-01-21
Training datasets
undisclosed
Added to our catalog
2026-07-28