Modelsintfloatmultilingual-e5-large-instruct
providerintfloat /

multilingual-e5-large-instruct

5 DZD in / 1M tokens

multilingual-e5-large-instruct is an open-source, instruction-tuned multilingual text embedding model developed by Microsoft Research. Built upon the 568-million parameter XLM-RoBERTa-large architecture, it supports over 100 languages and maps text into a 1024-dimensional dense vector space. Unlike previous E5 models that relied on rigid static prefixes, this model utilizes flexible, task-specific natural language instructions for queries while keeping target documents prefix-free, delivering superior custom retrieval accuracy and cross-lingual performance for global enterprise RAG systems.

Public
multilingual-e5-large-instruct
ArchitectureTransformer
Context Window512

1. Architectural & Technical Specifications

  • Developer / Organization: Microsoft Research (released under intfloat)
  • Base Model & Lineage: XLM-RoBERTa-large (E5 Instruction Series)
  • Release Date: Early 2024
  • Model Architecture: Encoder-only Transformer (Bidirectional)
  • Layer & Head Count: 24 Layers, 16 Attention heads
  • Total Parameters: ~568 Million
  • Active Parameters: ~568 Million
  • Embedding Dimension: 1024 dimensions
  • Vocabulary Size: ~250,002 tokens (SentencePiece)
  • Native Context Window: 512 tokens
  • Max Output Length: Fixed-length 1024-dimensional dense vector
  • License: MIT License (Permissive Open Source)

2. Training Data & Alignment Pipeline

  • Dataset Composition: Pre-trained on a massive cross-lingual dataset across 100+ languages, followed by fine-tuning on high-quality multilingual instruction-tuning datasets (covering retrieval, reranking, classification, and semantic textual similarity).
  • Fine-Tuning Methodology: Multi-task contrastive learning trained to adapt vector generation based on explicit task descriptions/instructions.
  • Instruction Support & Prompt Formatting:
    • Queries / Tasks: Formatted as Instruct: {task_description}\nQuery: {query}
    • Documents / Passages: Raw text directly without any prefix or instruction (unlike legacy E5).
    • Example Instruction: "Given a web search query, retrieve relevant passages that answer the query"
  • Similarity Metric: Cosine Similarity / Dot Product (requires L2 normalization).

3. Benchmark Performance & Statistics

Benchmark / MetricScore / ResultRelative Performance vs Base/Competitors
MTEB Multilingual Retrieval~65.8 (Avg)Outperforms standard multilingual-e5-large by 1.5-3%
MIRACL (Cross-Lingual Retrieval)~66.2 (nDCG@10)State-of-the-art among sub-billion parameter multilingual encoders
Multilingual Classification (MTEB)HighCustom task prompts boost classification separation significantly
Cross-Lingual Zero-ShotTop TierSeamless semantic alignment across diverse language pairs (e.g., AR/FR/EN)

4. Direct Comparative Analysis

Feature / Attributemultilingual-e5-large-instructmultilingual-e5-largebge-m3text-embedding-3-large
Parameter Count~568M~568M~567MClosed
Embedding Size1024 dim1024 dim1024 dim3072 dim
Context Window512 tokens512 tokens8,192 tokens8,191 tokens
Query FormattingTask Instruction (Instruct: ...\nQuery: ...)Fixed Prefix (query: )Optional PrefixNone
Passage FormattingNone (Raw text)Fixed Prefix (passage: )None (Raw text)None (Raw text)
Task CustomizationVery High (Promptable)Low (Fixed)ModerateModerate

5. Hardware Requirements & Quantization Specs

Deployment FormatFile Size (approx.)Minimum VRAM / RAMRecommended Target Device
FP32 (PyTorch/HF)~2.27 GB~3.5 GB RAM / VRAMCPU Server / Dedicated GPU (T4 / A10G)
FP16 (GPU Optimized)~1.14 GB~2 GB VRAMNVIDIA T4 / RTX 3060 / RTX 4090
ONNX Runtime (FP16/INT8)~570 MB - 1.14 GB~1.5 GB RAMHigh-throughput production microservices
INT8 / Quantized (ONNX/Q8)~570 MB~1 GB RAMServerless containers / Edge nodes

6. Recommended Inference Parameters & Usage

  • Pooling Method: Mean pooling (Average pooling of the last hidden state over attention-masked tokens).
  • Normalization: normalize_embeddings = True (Unit sphere L2 normalization for cosine similarity search).
  • Query Formulation:
    PYTHON
    def get_detailed_instruct(task_description: str, query: str) -> str:
        return f"Instruct: {task_description}\nQuery: {query}"