Modelsintfloate5-base-v2
providerintfloat /

e5-base-v2

3 DZD in / 1M tokens

e5-base-v2 is an open-source text embedding model developed by Microsoft Research. Based on a 109-million parameter BERT architecture, it is pre-trained using weakly-supervised contrastive learning. It generates 768-dimensional dense vectors and requires specific text prefixes ("query: " and "passage: ") to distinguish between search intents and document indexing, offering a strong balance between high retrieval accuracy and minimal computational cost for semantic search and RAG applications.

Public
e5-base-v2
ArchitectureTransformer
Context Window512

1. Architectural & Technical Specifications

  • Developer / Organization: Microsoft Research (released under intfloat)
  • Base Model & Lineage: bert-base-uncased (E5: Embeddings from bidirEctional Encoder rEpresentations)
  • Release Date: December 2022 (v2 updates in early 2023)
  • Model Architecture: Encoder-only Transformer (Bidirectional)
  • Layer & Head Count: 12 Layers, 12 Attention heads
  • Total Parameters: ~109 Million
  • Active Parameters: ~109 Million
  • Embedding Dimension: 768 dimensions
  • Vocabulary Size: ~30,522 tokens
  • Native Context Window: 512 tokens
  • Max Output Length: Fixed 768-dimensional dense vector
  • License: MIT License (Permissive Open Source)

2. Training Data & Alignment Pipeline

  • Dataset Composition: Pre-trained on CCPairs (a massive dataset of text pairs collected from web pages, Reddit, StackOverflow, etc.) consisting of hundreds of millions of diverse text pairs.
  • Fine-Tuning Methodology: Weakly-supervised contrastive pre-training followed by supervised fine-tuning on high-quality labeled datasets (like MS-MARCO and NLI).
  • System Prompt Adherence: Strict prefix dependency. Requires "query: " prepended to search queries, and "passage: " prepended to the documents being indexed.
  • Similarity Metric: Cosine Similarity (requires embeddings to be L2-normalized).

3. Benchmark Performance & Statistics

Benchmark / MetricScore / ResultRelative Performance vs Base/Competitors
MTEB Overall (English)~61.5Strong baseline, outperforms older models like all-mpnet-base-v2
MTEB Retrieval~50.2Solid performance for its size and age
MTEB Semantic Textual Similarity (STS)~82.4High correlation with human judgment
Latency / EfficiencyVery LowExcellent throughput on consumer hardware

4. Direct Comparative Analysis

Feature / Attributee5-base-v2e5-large-v2bge-base-en-v1.5all-MiniLM-L6-v2
Parameter Count~109M~335M~109M~22M
Embedding Size768 dim1024 dim768 dim384 dim
Context Length512 tokens512 tokens512 tokens256 / 512 tokens
Input PrefixesRequired (query: / passage:)RequiredRequired for queries onlyNot required
Overall MTEB~61.5~62.2~63.5~56.2

5. Hardware Requirements & Quantization Specs

Deployment FormatFile Size (approx.)Minimum VRAM / RAMRecommended Target Device
FP32 (PyTorch/HF)~438 MB~1 GB RAM / VRAMCPU (x86_64, ARM) / Entry Server GPU
FP16 (GPU Optimized)~219 MB~512 MB VRAMNVIDIA T4 / RTX 3060 / Edge Devices
ONNX Runtime (FP32/FP16)~220 - 438 MB~500 MB RAMProduction microservices / Serverless
INT8 / Quantized (ONNX/Q8)~110 MB~256 MB RAMMobile / Serverless / Low-memory nodes

6. Recommended Inference Parameters & Usage

  • Pooling Method: Mean pooling (Average pooling of the last hidden state).
  • Normalization: normalize_embeddings = True (Crucial for dot-product search equivalent to Cosine distance).
  • Query Prefix: "query: " (Must be added to the search string).
  • Passage Prefix: "passage: " (Must be added to the text before inserting into the vector database).
  • Framework Support: Fully supported in HuggingFace Transformers, Sentence-Transformers, and natively integrated into vector DB SDKs like PyMilvus.

7. Core Strengths & Deployment Recommendations

  • Key Strengths: Very fast and lightweight; highly reliable standard for general-purpose embedding tasks; extremely permissive MIT license.
  • Known Limitations: The strict requirement to prepend "query: " and "passage: " is often forgotten by developers, leading to degraded retrieval performance. The 512-token limit requires aggressive chunking. Eclipsed in pure accuracy by newer generation models like bge-base-en-v1.5.
  • Best Use Cases: Lightweight RAG pipelines, semantic search engines on tight compute budgets, and text clustering where processing speed and low API cost are prioritized over absolute state-of-the-art accuracy.