1,011 models · refreshed nightly

Audio classification models

Every model in the catalog with its licence, estimated VRAM and daily-tracked downloads. Filters update the URL — share any view.

#ModelParamsContextCommercial use30dMin VRAM
01 clap-htsat-fused laion · audio-classification 150 M ✓ apache-2.0 6.9 M from 1.2 GB
02 ast-finetuned-audioset-10-10-0.4593 MIT · audio-classification 90 M ✓ bsd-3-clause 1.2 M from 0.9 GB
03 wav2vec2-large-robust-12-ft-emotion-msp-dim audeering · audio-classification 170 M ✗ cc-by-nc-sa-4.0 751 K from 1.3 GB
04 wav2vec2-large-xlsr-53-gender-recognition-librispeech alefiury · audio-classification 320 M ✓ apache-2.0 626 K from 1.9 GB
05 MERT-v1-330M m-a-p · audio-classification ✗ cc-by-nc-4.0 560 K
06 audiobox-aesthetics facebook · audio-classification 100 M ✓ cc-by-4.0 401 K from 1.0 GB
07 voice-gender-classifier JaesungHuh · audio-classification 20 M ✓ mit 308 K from 0.6 GB
08 mms-lid-126 facebook · audio-classification 970 M ✗ cc-by-nc-4.0 308 K from 4.9 GB
09 MuQ-large-msd-iter OpenMuQ · audio-classification 330 M ✗ cc-by-nc-4.0 292 K from 2.0 GB
10 open-vakgyata onecxi · audio-classification 60 M ✗ cc-by-nc-4.0 278 K from 0.8 GB
11 wav2vec2-large-robust-24-ft-age-gender audeering · audio-classification 320 M ✗ cc-by-nc-sa-4.0 255 K from 1.9 GB
12 hubert-large-speech-emotion-recognition-russian-dusha-finetuned xbgoose · audio-classification 320 M ✓ apache-2.0 248 K from 1.9 GB
VRAM figures are estimates for the smallest available quantization at 8K context — see /methodology. Downloads refresh nightly from the Hugging Face API.