microsoft / text-classification updated 5 years ago

deberta-large-mnli

DeBERTa improves the BERT and RoBERTa models using disentangled attention and enhanced mask decoder. It outperforms BERT and RoBERTa on majority of NLU tasks with 80GB training data.

Params
Context
512
Downloads 30d
274 K
Likes
33
Commercial use: allowed mit Not gated 1 languages View on Hugging Face ↗

Download history

daily snapshots · 10 days
276 K266 K
Jul 28Jul 31Aug 3Aug 6

Specifications

Architecture
DebertaForSequenceClassification
Context length
512
Vocabulary
50,265
Layers / heads
24 / 16
Licence
mit
First seen on the Hub
2022-03-02
Training datasets
undisclosed
Added to our catalog
2026-07-28