mms-300m-1130-forced-aligner
This Python package provides an efficient way to perform forced alignment between text and audio using Hugging Face's pretrained models. it also features an improved implementation to use much less memory than TorchAudio forced alignment API.
Params
320 M
Context
—
Downloads 30d
2.7 M
Likes
103
Commercial use: not allowed · cc-by-nc-4.0
Not gated
SAFETENSORS
158 languages
View on Hugging Face ↗
Download history
daily snapshots · 55 days
▲ 198 K in the last 30 days (7.9%)
2.8 M2.4 M
Aug 22Sep 1Sep 11Sep 20
2.8 M2.1 M
Jul 28Aug 15Sep 2Sep 20
Can you run it?
Estimated VRAM at 8K context unless noted. Pick your hardware to see the verdict per quantization.
| File | Quant | Size | Est. VRAM | Verdict on RTX 4090 · 24 GB |
|---|---|---|---|---|
| model.safetensors | f32 | 1.3 GB | 1.9 GB | ✅ Runs comfortably |
Estimate: file size × 1.1 + KV cache at 8K + 0.5 GB overhead. Not a benchmark — how we calculate this.
Specifications
- Architecture
- Wav2Vec2ForCTC
- Parameters
- 320 M
- Tensor type
- F32
- Vocabulary
- 31
- Layers / heads
- 24 / 16
- Licence
- cc-by-nc-4.0
- First seen on the Hub
- 2024-05-02
- Training datasets
- undisclosed
- Added to our catalog
- 2026-07-28
Compare with any automatic-speech-recognition model