eddiegulay / automatic-speech-recognition updated 2 years ago

wav2vec2-large-xlsr-mvc-swahili

There was an issue with vocab, seems like there are special characters included and they were not considered during training You could try python from transformers import AutoProcessor, AutoModelForCTC

Params
320 M
Context
Downloads 30d
476 K
Likes
3
Commercial use: allowed apache-2.0 Not gated SAFETENSORS 1 languages View on Hugging Face ↗

Download history

daily snapshots · 10 days
610 K408 K
Jul 28Jul 31Aug 3Aug 6

Can you run it?

Estimated VRAM at 8K context unless noted. Pick your hardware to see the verdict per quantization.

FileQuantSizeEst. VRAMVerdict on RTX 4090 · 24 GB
model.safetensors f32 1.3 GB 1.9 GB ✅ Runs comfortably
Estimate: file size × 1.1 + KV cache at 8K + 0.5 GB overhead. Not a benchmark — how we calculate this.

Specifications

Architecture
Wav2Vec2ForCTC
Parameters
320 M
Tensor type
F32
Vocabulary
59
Layers / heads
24 / 16
Licence
apache-2.0
First seen on the Hub
2023-11-06
Training datasets
common_voice_13_0
common_voice_13_0 (reported)
0.2
Added to our catalog
2026-07-28