Modelsintfloate5-large-v2
providerintfloat /

e5-large-v2

5 DZD in / 1M tokens

e5-large-v2 is an open-source text embedding model developed by Microsoft Research. Based on a 335-million parameter BERT-large encoder architecture, it maps text into a dense 1024-dimensional vector space using weakly-supervised contrastive pre-training followed by supervised fine-tuning. It utilizes asymmetric prefix conditioning ("query: " and "passage: ") to optimize dense retrieval and semantic search, delivering strong zero-shot retrieval accuracy across diverse domains with predictable inference latency.

Public
e5-large-v2
ArchitectureTransformer
Context Window512

1. Architectural & Technical Specifications

  • Developer / Organization: Microsoft Research (released under intfloat)
  • Base Model & Lineage: bert-large-uncased (E5: Embeddings from bidirEctional Encoder rEpresentations)
  • Release Date: December 2022 (v2 updates in early 2023)
  • Model Architecture: Encoder-only Transformer (Bidirectional)
  • Layer & Head Count: 24 Layers, 16 Attention heads
  • Total Parameters: ~335 Million
  • Active Parameters: ~335 Million
  • Embedding Dimension: 1024 dimensions
  • Vocabulary Size: ~30,522 tokens (WordPiece)
  • Native Context Window: 512 tokens
  • Max Output Length: Fixed-length 1024-dimensional dense vector
  • License: MIT License (Permissive Open Source)

2. Training Data & Alignment Pipeline

  • Dataset Composition: Pre-trained on CCPairs (a curated web-scale corpus of ~270M diverse text pairs extracted from Reddit, StackExchange, Wikipedia, and Common Crawl).
  • Fine-Tuning Methodology: Multi-task contrastive learning combining weakly-supervised pre-training with supervised fine-tuning on high-quality labeled datasets (MS-MARCO, NLI, and QA datasets).
  • System Prompt / Prefix Conditioning: Asymmetric prefix requirement:
    • Search queries must be prefixed with "query: "
    • Documents/passages must be prefixed with "passage: "
  • Similarity Metric: Cosine Similarity / Dot Product (requires L2-normalized embeddings).

3. Benchmark Performance & Statistics

Benchmark / MetricScore / ResultRelative Performance vs Base/Competitors
MTEB Overall (English)~62.25Strong sub-billion parameter baseline
MTEB Retrieval (BEIR)~50.56Solid dense retrieval across diverse domains
MTEB Semantic Textual Similarity (STS)~83.11High alignment with human semantic rankings
MTEB Reranking~56.49Reliable zero-shot document ranking

4. Direct Comparative Analysis

Feature / Attributee5-large-v2e5-base-v2bge-large-en-v1.5gte-large
Parameter Count~335M~109M~335M~335M
Embedding Size1024 dim768 dim1024 dim1024 dim
Context Window512 tokens512 tokens512 tokens512 tokens
Prefix RequirementBoth (query: / passage:)Both (query: / passage:)Query onlyQuery only
MTEB Overall~62.25~61.50~64.11~63.13
Throughput / LatencyModerateVery FastModerateModerate

5. Hardware Requirements & Quantization Specs

Deployment FormatFile Size (approx.)Minimum VRAM / RAMRecommended Target Device
FP32 (PyTorch/HF)~1.34 GB~2 GB RAM / VRAMCPU Server / Dedicated GPU (T4 / A10G)
FP16 (GPU Optimized)~670 MB~1.5 GB VRAMNVIDIA T4 / RTX 3060 / RTX 4090
ONNX Runtime (FP32/FP16)~670 MB - 1.34 GB~1 GB RAMHigh-throughput production microservices
INT8 / Quantized (ONNX/Q8)~340 MB~512 MB RAMServerless containers / Edge nodes

6. Recommended Inference Parameters & Usage

  • Pooling Method: Mean pooling (Average pooling of the last hidden state over valid attention tokens).
  • Normalization: normalize_embeddings = True (Unit sphere L2 normalization).
  • Query Prefix: "query: " (Strictly mandatory for retrieval queries).
  • Passage Prefix: "passage: " (Strictly mandatory before vector database indexing).
  • Vector Database Integration: Native support in sentence-transformers, Hugging Face transformers, and compatible with pgvector, Qdrant, Milvus, and Pinecone.

7. Core Strengths & Deployment Recommendations

  • Key Strengths: Reliable and consistent retrieval quality, robust out-of-domain generalization, fast inference relative to LLM-based embedders, fully permissive MIT license for commercial applications.
  • Known Limitations: Mandatory "query: " and "passage: " prefixes create potential pipeline errors if omitted; 512-token context ceiling requires chunking long documents; slightly outperformed on MTEB benchmarks by newer generation models like bge-large-en-v1.5.
  • Best Use Cases: Production RAG systems, enterprise semantic search engines, automated document classification, and similarity-based clustering pipelines.