bosonai / text-to-speech updated 3 weeks ago

higgs-tts-3-4b

Higgs TTS 3 is built for voice chat: it speaks, not just reads. It turns model responses into expressive conversational speech across 100+ languages, with zero-shot voice cloning and inline control over emotion, style, prosody, pauses, and sound effects.

Params
4.7 B
Context
Downloads 30d
374 K
Likes
697
Licence: other Not gated SAFETENSORS 100 languages View on Hugging Face ↗

Download history

daily snapshots · 10 days
395 K374 K
Jul 28Jul 31Aug 3Aug 6

Can you run it?

Estimated VRAM at 8K context unless noted. Pick your hardware to see the verdict per quantization.

FileQuantSizeEst. VRAMVerdict on RTX 4090 · 24 GB
model.safetensors bf16 9.3 GB 11.4 GB ✅ Runs comfortably
Estimate: file size × 1.1 + KV cache at 8K + 0.5 GB overhead. Not a benchmark — how we calculate this.

Specifications

Architecture
HiggsMultimodalQwen3ForConditionalGeneration
Parameters
4.7 B
Tensor type
BF16
Licence
other
First seen on the Hub
2026-06-04
Training datasets
undisclosed
Added to our catalog
2026-07-28