ModelsBAAIbge-en-icl
providerBAAI /

bge-en-icl

5 DZD in / 1M tokens

bge-en-icl is a 7.11 billion parameter large language model-based embedding model developed by the Beijing Academy of Artificial Intelligence (BAAI), built upon the Mistral-7B backbone. It introduces powerful in-context learning (ICL) capabilities to text embedding generation, allowing users to provide task-specific examples (few-shot prompts) alongside their queries. This eliminates the need for fine-tuning for new tasks while achieving state-of-the-art retrieval and semantic representation performance on benchmarks like MTEB and AIR-Bench.

Public
bge-en-icl
ArchitectureTransformer
Context Window8K

1. Architectural & Technical Specifications

  • Developer / Organization: Beijing Academy of Artificial Intelligence (BAAI)
  • Base Model & Lineage: Mistral-7B (Transformer Decoder-Only Architecture)
  • Release Date: July 2024
  • Model Architecture: Decoder-Only, leveraging Mistral's Grouped-Query Attention (GQA) and Sliding Window Attention (SWA), adapted for embedding extraction (typically utilizing the last hidden state of a designated token).
  • Total Parameters: 7.11 Billion
  • Active Parameters: 7.11 Billion
  • Embedding Dimension: 4096 dimensions
  • Native Context Window: 8,192 tokens
  • License: MIT License (Permissive Open Source)

2. Training Data & Alignment Pipeline

  • Methodology: InfoNCE loss function; trained to maximize similarity between positive pairs while minimizing it for negative pairs.
  • In-Context Learning (ICL): Uniquely designed to accept task definitions and few-shot examples prepended directly to the input query, dynamically adjusting its output vector representation based on the provided context.
  • System Prompt / Instruction Adherence: Highly reliant on custom instructions. Example for retrieval: "Given a web search query, retrieve relevant passages that answer the query."
  • Flexibility: Capable of both zero-shot embedding and few-shot (ICL) embedding without any parameter updates.

3. Benchmark Performance & Statistics

Benchmark / MetricScore / ResultNote / Relative Performance
MTEB (Amazon Counterfactual)~93.15 (Accuracy/Main Score)Excellent classification capability
MTEB (Amazon Polarity)~96.98 (Accuracy/Main Score)Highly accurate binary classification
AIR-Bench (Zero-Shot)52.93 (Overall)Outperforms many 7B embedders without examples
AIR-Bench (Few-Shot)54.36 (Overall)State-of-the-Art performance; clear improvement via ICL

4. Direct Comparative Analysis

Feature / Attributebge-en-icl (BAAI)e5-mistral-7b-instructbge-large-en-v1.5
Base ArchitectureMistral-7BMistral-7BBERT-Large
Parameter Count~7B~7B~335M
Embedding Size4096 dim4096 dim1024 dim
In-Context LearningYes (Dynamic few-shot)Instruction-tuned (No ICL)No
Compute OverheadVery High (LLM-based)Very High (LLM-based)Low (Encoder-based)

5. Hardware Requirements & Quantization Specs

Deployment FormatFile Size (approx.)Minimum VRAM / RAMRecommended Target Device
FP32 / FP16 (HuggingFace)~28.5 GB (FP32)~16 GB VRAM1x 24GB GPU (e.g., RTX 3090/4090)
FP16 Inference~14-15 GB~16 GB VRAMHigh-end Consumer GPU / Server
Note: Due to being based on a 7B LLM, it requires significantly more compute and VRAM than traditional ~100M parameter encoder models.

6. Recommended Inference Parameters & Usage

  • Framework: Best utilized via FlagEmbedding library (FlagICLModel) or HuggingFace Transformers.
  • Instruction Format: Requires task-specific query instructions (e.g., "Given a web search query, retrieve relevant passages that answer the query.").
  • ICL Examples Formulation: Provide examples as JSON-like dicts containing instruct, query, and response.
  • FP16 Execution: Recommended (use_fp16=True) to speed up computation with negligible performance degradation.

7. Core Strengths & Deployment Recommendations

  • Key Strengths: Unprecedented adaptability to niche, domain-specific tasks using only a few examples in the prompt; exceptional baseline performance even in zero-shot scenarios.
  • Known Limitations: Substantially higher latency and infrastructure costs compared to traditional BERT-based embedders (e.g., bge-base-en-v1.5); produces large 4096-dimensional vectors which increase vector database storage costs.
  • Best Use Cases: High-end RAG systems requiring extreme accuracy, complex classification or clustering tasks, and scenarios where the embedding task frequently changes or is highly specialized without resources for fine-tuning.