sumedh / automatic-speech-recognition updated 1 year ago

wav2vec2-large-xlsr-marathi

Fine-tuned facebook/wav2vec2-large-xlsr-53 on Marathi using the Open SLR64 dataset. When using this model, make sure that your speech input is sampled at 16kHz. This data contains only female voices but the model works well for male voices too. Trained on Google Colab Pro on Tesla P100 16GB GPU. WER (Word Error Rate) on the Test Set: 12.70 % The model can be used directly without a language model as follows, given th...

Params
320 M
Context
Downloads 30d
399 K
Likes
2
Commercial use: allowed apache-2.0 Not gated SAFETENSORS 1 languages View on Hugging Face ↗

Download history

daily snapshots · 25 days
489 K397 K
Aug 28Sep 5Sep 13Sep 21

Can you run it?

Estimated VRAM at 8K context unless noted. Pick your hardware to see the verdict per quantization.

FileQuantSizeEst. VRAMVerdict on RTX 4090 · 24 GB
model.safetensors f32 1.3 GB 1.9 GB ✅ Runs comfortably
Estimate: file size × 1.1 + KV cache at 8K + 0.5 GB overhead. Not a benchmark — how we calculate this.

Specifications

Architecture
Wav2Vec2ForCTC
Parameters
320 M
Tensor type
F32
Vocabulary
68
Layers / heads
24 / 16
Licence
apache-2.0
First seen on the Hub
2022-03-02
Training datasets
openslr
OpenSLR mr (reported)
12.7
Added to our catalog
2026-08-28
Compare with any automatic-speech-recognition model