1,054 models · refreshed nightly

Audio classification models

Every model in the catalog with its licence, estimated VRAM and daily-tracked downloads. Filters update the URL — share any view.

#ModelParamsContextCommercial use30dMin VRAM
01 clap-htsat-fused laion · audio-classification 150 M ✓ apache-2.0 8.2 M from 1.2 GB
02 wav2vec2-large-robust-24-ft-age-gender audeering · audio-classification 320 M ✗ cc-by-nc-sa-4.0 2.5 M from 1.9 GB
03 ced-gguf mudler · audio-classification ✓ apache-2.0 922 K from 0.6 GB
04 wav2vec2-large-robust-12-ft-emotion-msp-dim audeering · audio-classification 170 M ✗ cc-by-nc-sa-4.0 744 K from 1.3 GB
05 wav2vec2-large-xlsr-53-gender-recognition-librispeech alefiury · audio-classification 320 M ✓ apache-2.0 714 K from 1.9 GB
06 ast-finetuned-audioset-10-10-0.4593 MIT · audio-classification 90 M ✓ bsd-3-clause 705 K from 0.9 GB
07 audiobox-aesthetics facebook · audio-classification 100 M ✓ cc-by-4.0 641 K from 1.0 GB
08 open-vakgyata onecxi · audio-classification 60 M ✗ cc-by-nc-4.0 504 K from 0.8 GB
09 MuQ-large-msd-iter OpenMuQ · audio-classification 330 M ✗ cc-by-nc-4.0 295 K from 2.0 GB
10 voice-gender-classifier JaesungHuh · audio-classification 20 M ✓ mit 291 K from 0.6 GB
11 hubert-large-speech-emotion-recognition-russian-dusha-finetuned xbgoose · audio-classification 320 M ✓ apache-2.0 219 K from 1.9 GB
VRAM figures are estimates for the smallest available quantization at 8K context — see /methodology. Downloads refresh nightly from the Hugging Face API.