clipseg-rd64-refined vs mask2former-swin-large-cityscapes-semantic
Specs, VRAM requirements and download trends — updated 21 September 2026.
Specification comparison
Differences are highlighted; identical values are muted.
| Specification | clipseg-rd64-refined | mask2former-swin-large-cityscapes-semantic |
|---|---|---|
| Parameters | 150 M | 220 M |
| Architecture | CLIPSegForImageSegmentation | Mask2FormerForUniversalSegmentation |
| Context length | — | — |
| Licence | apache-2.0 | other |
| Commercial use | Allowed | Unknown |
| Languages | — | — |
| Downloads 30d | 1,096,224 | 444,440 |
| Downloads all time | 152.2 M | 10.6 M |
| Quantizations on the Hub | SAFETENSORS | SAFETENSORS |
| Gated | No | No |
| First seen on the Hub | 2022-11-01 | 2023-01-05 |
Download trend
Daily snapshots, last 56 days (28 Jul – 21 Sep)
1.3 M239 K
clipseg-rd64-refined
mask2former-swin-large-cityscapes-semantic
VRAM side by side
On RTX 4090 · 24 GB · 8K context unless noted
clipseg-rd64-refined · fp16
1.2 GB / 24 GB
mask2former-swin-large-cityscapes-semantic · fp16
1.5 GB / 24 GB