VibeVoice-Realtime-0.5B
VibeVoice-Realtime is a lightweight real‑time text-to-speech model supporting streaming text input and robust long-form speech generation. It can be used to build realtime TTS services, narrate live data streams, and let different LLMs start speaking from their very first tokens (plug in your preferred model) long before a full answer is generated. It produces initial audible speech in ~300 ms (hardware dependent).
Params
1.0 B
Context
—
Downloads 30d
657 K
Likes
1,263
Download history
daily snapshots · 10 days662 K650 K
Jul 28Jul 31Aug 3Aug 6
Can you run it?
Estimated VRAM at 8K context unless noted. Pick your hardware to see the verdict per quantization.
| File | Quant | Size | Est. VRAM | Verdict on RTX 4090 · 24 GB |
|---|---|---|---|---|
| model.safetensors | bf16 | 2.0 GB | 2.9 GB | ✅ Runs comfortably |
Estimate: file size × 1.1 + KV cache at 8K + 0.5 GB overhead. Not a benchmark — how we calculate this.
Specifications
- Architecture
- VibeVoiceStreamingForConditionalGenerationInference
- Parameters
- 1.0 B
- Tensor type
- BF16
- Licence
- mit
- First seen on the Hub
- 2025-12-04
- Base model
- Qwen2.5-0.5B
- Training datasets
- undisclosed
- Added to our catalog
- 2026-07-28
Compare with
Sponsored · GPU cloud
Not enough VRAM?
Spin up a 24 GB L4 instance in 40 seconds. $0.44/hr.