facebook / image-classification updated 1 year ago

convnextv2-base-22k-384

ConvNeXt V2 model pretrained using the FCMAE framework and fine-tuned on the ImageNet-22K dataset at resolution 384x384. It was introduced in the paper ConvNeXt V2: Co-designing and Scaling ConvNets with Masked Autoencoders by Woo et al. and first released in this repository.

Params
90 M
Context
Downloads 30d
259 K
Likes
4
Commercial use: allowed apache-2.0 Not gated SAFETENSORS View on Hugging Face ↗

Download history

daily snapshots · 44 days
▲ 202 K in the last 30 days (43.8%)
476 K244 K
Aug 8Aug 22Sep 6Sep 20

Can you run it?

Estimated VRAM at 8K context unless noted. Pick your hardware to see the verdict per quantization.

FileQuantSizeEst. VRAMVerdict on RTX 4090 · 24 GB
model.safetensors f32 0.4 GB 0.9 GB ✅ Runs comfortably
Estimate: file size × 1.1 + KV cache at 8K + 0.5 GB overhead. Not a benchmark — how we calculate this.

Specifications

Architecture
ConvNextV2ForImageClassification
Parameters
90 M
Tensor type
F32
Licence
apache-2.0
First seen on the Hub
2023-02-19
Training datasets
imagenet-22k
Added to our catalog
2026-08-08
Compare with any image-classification model