audeering / audio-classification updated 1 year ago

wav2vec2-large-robust-24-ft-age-gender

The model expects a raw audio signal as input and outputs predictions for age in a range of approximately 0...1 (0...100 years) and gender expressing the probababilty for being child, female, or male. In addition, it also provides the pooled states of the last transformer layer. The model was created by fine-tuning Wav2Vec2-Large-Robust on aGender, Mozilla Common Voice, Timit and Voxceleb 2. For this version of the m...

Params
320 M
Context
Downloads 30d
255 K
Likes
57
Commercial use: not allowed · cc-by-nc-sa-4.0 Not gated SAFETENSORS View on Hugging Face ↗

Download history

daily snapshots · 10 days
276 K255 K
Jul 28Jul 31Aug 3Aug 6

Can you run it?

Estimated VRAM at 8K context unless noted. Pick your hardware to see the verdict per quantization.

FileQuantSizeEst. VRAMVerdict on RTX 4090 · 24 GB
model.safetensors f32 1.3 GB 1.9 GB ✅ Runs comfortably
Estimate: file size × 1.1 + KV cache at 8K + 0.5 GB overhead. Not a benchmark — how we calculate this.

Specifications

Architecture
Model
Parameters
320 M
Tensor type
F32
Vocabulary
32
Layers / heads
24 / 16
Licence
cc-by-nc-sa-4.0
First seen on the Hub
2023-09-04
Training datasets
agender, mozillacommonvoice, timit, voxceleb2
Added to our catalog
2026-07-28