wav2vec2-large-xlsr-marathi
Fine-tuned facebook/wav2vec2-large-xlsr-53 on Marathi using the Open SLR64 dataset. When using this model, make sure that your speech input is sampled at 16kHz. This data contains only female voices but the model works well for male voices too. Trained on Google Colab Pro on Tesla P100 16GB GPU. WER (Word Error Rate) on the Test Set: 12.70 % The model can be used directly without a language model as follows, given th...
Params
320 M
Context
—
Downloads 30d
399 K
Likes
2
Download history
daily snapshots · 25 days489 K397 K
Aug 28Sep 5Sep 13Sep 21
Can you run it?
Estimated VRAM at 8K context unless noted. Pick your hardware to see the verdict per quantization.
| File | Quant | Size | Est. VRAM | Verdict on RTX 4090 · 24 GB |
|---|---|---|---|---|
| model.safetensors | f32 | 1.3 GB | 1.9 GB | ✅ Runs comfortably |
Estimate: file size × 1.1 + KV cache at 8K + 0.5 GB overhead. Not a benchmark — how we calculate this.
Specifications
- Architecture
- Wav2Vec2ForCTC
- Parameters
- 320 M
- Tensor type
- F32
- Vocabulary
- 68
- Layers / heads
- 24 / 16
- Licence
- apache-2.0
- First seen on the Hub
- 2022-03-02
- Training datasets
- openslr
- OpenSLR mr (reported)
- 12.7
- Added to our catalog
- 2026-08-28
Compare with any automatic-speech-recognition model