wav2vec2-large-robust-24-ft-age-gender
The model expects a raw audio signal as input and outputs predictions for age in a range of approximately 0...1 (0...100 years) and gender expressing the probababilty for being child, female, or male. In addition, it also provides the pooled states of the last transformer layer. The model was created by fine-tuning Wav2Vec2-Large-Robust on aGender, Mozilla Common Voice, Timit and Voxceleb 2. For this version of the m...
Params
320 M
Context
—
Downloads 30d
2.5 M
Likes
61
Download history
daily snapshots · 55 days
▲ 1.5 M in the last 30 days (146.5%)
2.5 M1.1 M
Aug 22Sep 1Sep 11Sep 20
2.5 M229 K
Jul 28Aug 15Sep 2Sep 20
Can you run it?
Estimated VRAM at 8K context unless noted. Pick your hardware to see the verdict per quantization.
| File | Quant | Size | Est. VRAM | Verdict on RTX 4090 · 24 GB |
|---|---|---|---|---|
| model.safetensors | f32 | 1.3 GB | 1.9 GB | ✅ Runs comfortably |
Estimate: file size × 1.1 + KV cache at 8K + 0.5 GB overhead. Not a benchmark — how we calculate this.
Specifications
- Architecture
- Model
- Parameters
- 320 M
- Tensor type
- F32
- Vocabulary
- 32
- Layers / heads
- 24 / 16
- Licence
- cc-by-nc-sa-4.0
- First seen on the Hub
- 2023-09-04
- Training datasets
- agender, mozillacommonvoice, timit, voxceleb2
- Added to our catalog
- 2026-07-28
Compare with any audio-classification model