MahmoudAshraf / automatic-speech-recognition updated 3 months ago

mms-300m-1130-forced-aligner

This Python package provides an efficient way to perform forced alignment between text and audio using Hugging Face's pretrained models. it also features an improved implementation to use much less memory than TorchAudio forced alignment API.

Params
320 M
Context
Downloads 30d
2.5 M
Likes
96
Commercial use: not allowed · cc-by-nc-4.0 Not gated SAFETENSORS 158 languages View on Hugging Face ↗

Download history

daily snapshots · 10 days
2.5 M2.4 M
Jul 28Jul 31Aug 3Aug 6

Can you run it?

Estimated VRAM at 8K context unless noted. Pick your hardware to see the verdict per quantization.

FileQuantSizeEst. VRAMVerdict on RTX 4090 · 24 GB
model.safetensors f32 1.3 GB 1.9 GB ✅ Runs comfortably
Estimate: file size × 1.1 + KV cache at 8K + 0.5 GB overhead. Not a benchmark — how we calculate this.

Specifications

Architecture
Wav2Vec2ForCTC
Parameters
320 M
Tensor type
F32
Vocabulary
31
Layers / heads
24 / 16
Licence
cc-by-nc-4.0
First seen on the Hub
2024-05-02
Training datasets
undisclosed
Added to our catalog
2026-07-28