e5-base-v2
3 DZD in / 1M tokens
e5-base-v2 is an open-source text embedding model developed by Microsoft Research. Based on a 109-million parameter BERT architecture, it is pre-trained using weakly-supervised contrastive learning. It generates 768-dimensional dense vectors and requires specific text prefixes ("query: " and "passage: ") to distinguish between search intents and document indexing, offering a strong balance between high retrieval accuracy and minimal computational cost for semantic search and RAG applications.
Public

ArchitectureTransformer
Context Window512
1. Architectural & Technical Specifications
- Developer / Organization: Microsoft Research (released under
intfloat) - Base Model & Lineage:
bert-base-uncased(E5: Embeddings from bidirEctional Encoder rEpresentations) - Release Date: December 2022 (v2 updates in early 2023)
- Model Architecture: Encoder-only Transformer (Bidirectional)
- Layer & Head Count: 12 Layers, 12 Attention heads
- Total Parameters: ~109 Million
- Active Parameters: ~109 Million
- Embedding Dimension: 768 dimensions
- Vocabulary Size: ~30,522 tokens
- Native Context Window: 512 tokens
- Max Output Length: Fixed 768-dimensional dense vector
- License: MIT License (Permissive Open Source)
2. Training Data & Alignment Pipeline
- Dataset Composition: Pre-trained on CCPairs (a massive dataset of text pairs collected from web pages, Reddit, StackOverflow, etc.) consisting of hundreds of millions of diverse text pairs.
- Fine-Tuning Methodology: Weakly-supervised contrastive pre-training followed by supervised fine-tuning on high-quality labeled datasets (like MS-MARCO and NLI).
- System Prompt Adherence: Strict prefix dependency. Requires
"query: "prepended to search queries, and"passage: "prepended to the documents being indexed. - Similarity Metric: Cosine Similarity (requires embeddings to be L2-normalized).
3. Benchmark Performance & Statistics
| Benchmark / Metric | Score / Result | Relative Performance vs Base/Competitors |
|---|---|---|
| MTEB Overall (English) | ~61.5 | Strong baseline, outperforms older models like all-mpnet-base-v2 |
| MTEB Retrieval | ~50.2 | Solid performance for its size and age |
| MTEB Semantic Textual Similarity (STS) | ~82.4 | High correlation with human judgment |
| Latency / Efficiency | Very Low | Excellent throughput on consumer hardware |
4. Direct Comparative Analysis
| Feature / Attribute | e5-base-v2 | e5-large-v2 | bge-base-en-v1.5 | all-MiniLM-L6-v2 |
|---|---|---|---|---|
| Parameter Count | ~109M | ~335M | ~109M | ~22M |
| Embedding Size | 768 dim | 1024 dim | 768 dim | 384 dim |
| Context Length | 512 tokens | 512 tokens | 512 tokens | 256 / 512 tokens |
| Input Prefixes | Required (query: / passage:) | Required | Required for queries only | Not required |
| Overall MTEB | ~61.5 | ~62.2 | ~63.5 | ~56.2 |
5. Hardware Requirements & Quantization Specs
| Deployment Format | File Size (approx.) | Minimum VRAM / RAM | Recommended Target Device |
|---|---|---|---|
| FP32 (PyTorch/HF) | ~438 MB | ~1 GB RAM / VRAM | CPU (x86_64, ARM) / Entry Server GPU |
| FP16 (GPU Optimized) | ~219 MB | ~512 MB VRAM | NVIDIA T4 / RTX 3060 / Edge Devices |
| ONNX Runtime (FP32/FP16) | ~220 - 438 MB | ~500 MB RAM | Production microservices / Serverless |
| INT8 / Quantized (ONNX/Q8) | ~110 MB | ~256 MB RAM | Mobile / Serverless / Low-memory nodes |
6. Recommended Inference Parameters & Usage
- Pooling Method: Mean pooling (Average pooling of the last hidden state).
- Normalization:
normalize_embeddings = True(Crucial for dot-product search equivalent to Cosine distance). - Query Prefix:
"query: "(Must be added to the search string). - Passage Prefix:
"passage: "(Must be added to the text before inserting into the vector database). - Framework Support: Fully supported in HuggingFace
Transformers,Sentence-Transformers, and natively integrated into vector DB SDKs like PyMilvus.
7. Core Strengths & Deployment Recommendations
- Key Strengths: Very fast and lightweight; highly reliable standard for general-purpose embedding tasks; extremely permissive MIT license.
- Known Limitations: The strict requirement to prepend
"query: "and"passage: "is often forgotten by developers, leading to degraded retrieval performance. The 512-token limit requires aggressive chunking. Eclipsed in pure accuracy by newer generation models likebge-base-en-v1.5. - Best Use Cases: Lightweight RAG pipelines, semantic search engines on tight compute budgets, and text clustering where processing speed and low API cost are prioritized over absolute state-of-the-art accuracy.