microsoft / text-classification updated 5 years ago

deberta-large-mnli

DeBERTa improves the BERT and RoBERTa models using disentangled attention and enhanced mask decoder. It outperforms BERT and RoBERTa on majority of NLU tasks with 80GB training data.

Params
Context
512
Downloads 30d
252 K
Likes
33
Commercial use: allowed mit Not gated 1 languages View on Hugging Face ↗

Download history

daily snapshots · 55 days
▲ 0 in the last 30 days (0.0%)
276 K233 K
Jul 28Aug 15Sep 2Sep 20

Specifications

Architecture
DebertaForSequenceClassification
Context length
512
Vocabulary
50,265
Layers / heads
24 / 16
Licence
mit
First seen on the Hub
2022-03-02
Training datasets
undisclosed
Added to our catalog
2026-07-28
Compare with any text-classification model