pnnbao-ump / text-to-speech updated 4 days ago

VieNeu-TTS-v3-Turbo

VieNeu-TTS v3 Turbo is the next generation of Vietnamese TTS — 48 kHz high-fidelity speech, 23 built-in preset voices across three regions (North / Central / South), instant voice cloning, real-time streaming with an OpenAI-compatible API (16 concurrent streams on one RTX 3060), inline emotion cues, and seamless bilingual (En–Vi) code-switching.

Params
130 M
Context
1,024
Downloads 30d
601 K
Likes
64
Commercial use: allowed apache-2.0 Not gated SAFETENSORS 2 languages View on Hugging Face ↗

Download history

daily snapshots · 52 days
▲ 246 K in the last 30 days (69.2%)
601 K306 K
Jul 31Aug 17Sep 3Sep 20

Can you run it?

Estimated VRAM at 8K context unless noted. Pick your hardware to see the verdict per quantization.

FileQuantSizeEst. VRAMVerdict on RTX 4090 · 24 GB
model.safetensors bf16 0.5 GB 1.1 GB ✅ Runs comfortably
Estimate: file size × 1.1 + KV cache at 8K + 0.5 GB overhead. Not a benchmark — how we calculate this.

Specifications

Architecture
VieNeuV3TurboForTTS
Parameters
130 M
Tensor type
BF16
Context length
1,024
Layers / heads
12 / 12
Licence
apache-2.0
First seen on the Hub
2026-06-05
Training datasets
pnnbao-ump/VieNeu-TTS-10k-ENVI
Added to our catalog
2026-07-31
Compare with any text-to-speech model