e5-large-v2
5 DZD in / 1M tokens
e5-large-v2 is an open-source text embedding model developed by Microsoft Research. Based on a 335-million parameter BERT-large encoder architecture, it maps text into a dense 1024-dimensional vector space using weakly-supervised contrastive pre-training followed by supervised fine-tuning. It utilizes asymmetric prefix conditioning ("query: " and "passage: ") to optimize dense retrieval and semantic search, delivering strong zero-shot retrieval accuracy across diverse domains with predictable inference latency.
Public

ArchitectureTransformer
Context Window512
1. Architectural & Technical Specifications
- Developer / Organization: Microsoft Research (released under
intfloat) - Base Model & Lineage:
bert-large-uncased(E5: Embeddings from bidirEctional Encoder rEpresentations) - Release Date: December 2022 (v2 updates in early 2023)
- Model Architecture: Encoder-only Transformer (Bidirectional)
- Layer & Head Count: 24 Layers, 16 Attention heads
- Total Parameters: ~335 Million
- Active Parameters: ~335 Million
- Embedding Dimension: 1024 dimensions
- Vocabulary Size: ~30,522 tokens (WordPiece)
- Native Context Window: 512 tokens
- Max Output Length: Fixed-length 1024-dimensional dense vector
- License: MIT License (Permissive Open Source)
2. Training Data & Alignment Pipeline
- Dataset Composition: Pre-trained on CCPairs (a curated web-scale corpus of ~270M diverse text pairs extracted from Reddit, StackExchange, Wikipedia, and Common Crawl).
- Fine-Tuning Methodology: Multi-task contrastive learning combining weakly-supervised pre-training with supervised fine-tuning on high-quality labeled datasets (MS-MARCO, NLI, and QA datasets).
- System Prompt / Prefix Conditioning: Asymmetric prefix requirement:
- Search queries must be prefixed with
"query: " - Documents/passages must be prefixed with
"passage: "
- Search queries must be prefixed with
- Similarity Metric: Cosine Similarity / Dot Product (requires L2-normalized embeddings).
3. Benchmark Performance & Statistics
| Benchmark / Metric | Score / Result | Relative Performance vs Base/Competitors |
|---|---|---|
| MTEB Overall (English) | ~62.25 | Strong sub-billion parameter baseline |
| MTEB Retrieval (BEIR) | ~50.56 | Solid dense retrieval across diverse domains |
| MTEB Semantic Textual Similarity (STS) | ~83.11 | High alignment with human semantic rankings |
| MTEB Reranking | ~56.49 | Reliable zero-shot document ranking |
4. Direct Comparative Analysis
| Feature / Attribute | e5-large-v2 | e5-base-v2 | bge-large-en-v1.5 | gte-large |
|---|---|---|---|---|
| Parameter Count | ~335M | ~109M | ~335M | ~335M |
| Embedding Size | 1024 dim | 768 dim | 1024 dim | 1024 dim |
| Context Window | 512 tokens | 512 tokens | 512 tokens | 512 tokens |
| Prefix Requirement | Both (query: / passage:) | Both (query: / passage:) | Query only | Query only |
| MTEB Overall | ~62.25 | ~61.50 | ~64.11 | ~63.13 |
| Throughput / Latency | Moderate | Very Fast | Moderate | Moderate |
5. Hardware Requirements & Quantization Specs
| Deployment Format | File Size (approx.) | Minimum VRAM / RAM | Recommended Target Device |
|---|---|---|---|
| FP32 (PyTorch/HF) | ~1.34 GB | ~2 GB RAM / VRAM | CPU Server / Dedicated GPU (T4 / A10G) |
| FP16 (GPU Optimized) | ~670 MB | ~1.5 GB VRAM | NVIDIA T4 / RTX 3060 / RTX 4090 |
| ONNX Runtime (FP32/FP16) | ~670 MB - 1.34 GB | ~1 GB RAM | High-throughput production microservices |
| INT8 / Quantized (ONNX/Q8) | ~340 MB | ~512 MB RAM | Serverless containers / Edge nodes |
6. Recommended Inference Parameters & Usage
- Pooling Method: Mean pooling (Average pooling of the last hidden state over valid attention tokens).
- Normalization:
normalize_embeddings = True(Unit sphere L2 normalization). - Query Prefix:
"query: "(Strictly mandatory for retrieval queries). - Passage Prefix:
"passage: "(Strictly mandatory before vector database indexing). - Vector Database Integration: Native support in
sentence-transformers, Hugging Facetransformers, and compatible with pgvector, Qdrant, Milvus, and Pinecone.
7. Core Strengths & Deployment Recommendations
- Key Strengths: Reliable and consistent retrieval quality, robust out-of-domain generalization, fast inference relative to LLM-based embedders, fully permissive MIT license for commercial applications.
- Known Limitations: Mandatory
"query: "and"passage: "prefixes create potential pipeline errors if omitted; 512-token context ceiling requires chunking long documents; slightly outperformed on MTEB benchmarks by newer generation models likebge-large-en-v1.5. - Best Use Cases: Production RAG systems, enterprise semantic search engines, automated document classification, and similarity-based clustering pipelines.