wav2vec2-large-robust-24-ft-age-gender
The model expects a raw audio signal as input and outputs predictions for age in a range of approximately 0...1 (0...100 years) and gender expressing the probababilty for being child, female, or male. In addition, it also provides the pooled states of the last transformer layer. The model was created by fine-tuning Wav2Vec2-Large-Robust on aGender, Mozilla Common Voice, Timit and Voxceleb 2. For this version of the m...
Params
320 M
Context
—
Downloads 30d
255 K
Likes
57
Download history
daily snapshots · 10 days276 K255 K
Jul 28Jul 31Aug 3Aug 6
Can you run it?
Estimated VRAM at 8K context unless noted. Pick your hardware to see the verdict per quantization.
| File | Quant | Size | Est. VRAM | Verdict on RTX 4090 · 24 GB |
|---|---|---|---|---|
| model.safetensors | f32 | 1.3 GB | 1.9 GB | ✅ Runs comfortably |
Estimate: file size × 1.1 + KV cache at 8K + 0.5 GB overhead. Not a benchmark — how we calculate this.
Specifications
- Architecture
- Model
- Parameters
- 320 M
- Tensor type
- F32
- Vocabulary
- 32
- Layers / heads
- 24 / 16
- Licence
- cc-by-nc-sa-4.0
- First seen on the Hub
- 2023-09-04
- Training datasets
- agender, mozillacommonvoice, timit, voxceleb2
- Added to our catalog
- 2026-07-28
Compare with
Sponsored · GPU cloud
Not enough VRAM?
Spin up a 24 GB L4 instance in 40 seconds. $0.44/hr.