BAAI / feature-extraction updated 2 years ago

bge-large-en-v1.5

If you are looking for a model that supports more languages, longer texts, and other retrieval methods, you can try using bge-m3.

Params
340 M
Context
512
Downloads 30d
13.1 M
Likes
708
Commercial use: allowed mit Not gated SAFETENSORS 1 languages View on Hugging Face ↗

Download history

daily snapshots · 10 days
13.4 M12.9 M
Jul 28Jul 31Aug 3Aug 6

Can you run it?

Estimated VRAM at 8K context unless noted. Pick your hardware to see the verdict per quantization.

FileQuantSizeEst. VRAMVerdict on RTX 4090 · 24 GB
model.safetensors f32 1.3 GB 2.0 GB ✅ Runs comfortably
Estimate: file size × 1.1 + KV cache at 8K + 0.5 GB overhead. Not a benchmark — how we calculate this.

Specifications

Architecture
BertModel
Parameters
340 M
Tensor type
F32
Context length
512
Vocabulary
30,522
Layers / heads
24 / 16
Licence
mit
First seen on the Hub
2023-09-12
Training datasets
undisclosed
MTEB ArguAna (reported)
77.027
MTEB BIOSSES (reported)
84.206737984757
MTEB ArxivClusteringP2P (reported)
48.56707792675
MTEB ArxivClusteringS2S (reported)
43.194533891824
MTEB BiorxivClusteringP2P (reported)
39.71453833765
MTEB BiorxivClusteringS2S (reported)
36.901083492844
MTEB AskUbuntuDupQuestions (reported)
77.823616057688
MTEB Banking77Classification (reported)
87.771893109649
MTEB CQADupstackAndroidRetrieval (reported)
32.795
MTEB AmazonPolarityClassification (reported)
92.394770195742
MTEB AmazonReviewsClassification (en) (reported)
47.807127928703
MTEB AmazonCounterfactualClassification (en) (reported)
69.693866480435
Added to our catalog
2026-07-28