Alibaba-NLP / sentence-similarity updated 1 year ago

gte-multilingual-base

The gte-multilingual-base model is the latest in the GTE (General Text Embedding) family of models, featuring several key attributes:

Params
310 M
Context
8,192
Downloads 30d
1.2 M
Likes
370
Commercial use: allowed apache-2.0 Not gated SAFETENSORS 75 languages View on Hugging Face ↗

Download history

daily snapshots · 10 days
1.2 M1.2 M
Jul 28Jul 31Aug 3Aug 6

Can you run it?

Estimated VRAM at 8K context unless noted. Pick your hardware to see the verdict per quantization.

FileQuantSizeEst. VRAMVerdict on RTX 4090 · 24 GB
model.safetensors f16 0.6 GB 1.2 GB ✅ Runs comfortably
Estimate: file size × 1.1 + KV cache at 8K + 0.5 GB overhead. Not a benchmark — how we calculate this.

Specifications

Architecture
NewModel
Parameters
310 M
Tensor type
F16
Context length
8,192
Vocabulary
250,048
Layers / heads
12 / 12
Licence
apache-2.0
First seen on the Hub
2024-07-20
Training datasets
undisclosed
MTEB ATEC (reported)
48.911863634175
MTEB AFQMC (reported)
43.54760696384
MTEB AllegroReviews (reported)
41.68986083499
MTEB 8TagsClustering (reported)
33.6668172633
MTEB AlloprofReranking (reported)
64.91495250072
MTEB AlloprofRetrieval (reported)
53.638
MTEB AlloProfClusteringP2P (reported)
54.202413379779
MTEB AlloProfClusteringS2S (reported)
44.340836956086
MTEB AmazonPolarityClassification (reported)
80.717625
MTEB AmazonReviewsClassification (de) (reported)
40.108
MTEB AmazonReviewsClassification (en) (reported)
43.642
MTEB AmazonCounterfactualClassification (en) (reported)
75.955223880597
Added to our catalog
2026-07-28