bge-en-icl
5 DZD in / 1M tokens
bge-en-icl is a 7.11 billion parameter large language model-based embedding model developed by the Beijing Academy of Artificial Intelligence (BAAI), built upon the Mistral-7B backbone. It introduces powerful in-context learning (ICL) capabilities to text embedding generation, allowing users to provide task-specific examples (few-shot prompts) alongside their queries. This eliminates the need for fine-tuning for new tasks while achieving state-of-the-art retrieval and semantic representation performance on benchmarks like MTEB and AIR-Bench.
Public

ArchitectureTransformer
Context Window8K
1. Architectural & Technical Specifications
- Developer / Organization: Beijing Academy of Artificial Intelligence (BAAI)
- Base Model & Lineage: Mistral-7B (Transformer Decoder-Only Architecture)
- Release Date: July 2024
- Model Architecture: Decoder-Only, leveraging Mistral's Grouped-Query Attention (GQA) and Sliding Window Attention (SWA), adapted for embedding extraction (typically utilizing the last hidden state of a designated token).
- Total Parameters: 7.11 Billion
- Active Parameters: 7.11 Billion
- Embedding Dimension: 4096 dimensions
- Native Context Window: 8,192 tokens
- License: MIT License (Permissive Open Source)
2. Training Data & Alignment Pipeline
- Methodology: InfoNCE loss function; trained to maximize similarity between positive pairs while minimizing it for negative pairs.
- In-Context Learning (ICL): Uniquely designed to accept task definitions and few-shot examples prepended directly to the input query, dynamically adjusting its output vector representation based on the provided context.
- System Prompt / Instruction Adherence: Highly reliant on custom instructions. Example for retrieval:
"Given a web search query, retrieve relevant passages that answer the query." - Flexibility: Capable of both zero-shot embedding and few-shot (ICL) embedding without any parameter updates.
3. Benchmark Performance & Statistics
| Benchmark / Metric | Score / Result | Note / Relative Performance |
|---|---|---|
| MTEB (Amazon Counterfactual) | ~93.15 (Accuracy/Main Score) | Excellent classification capability |
| MTEB (Amazon Polarity) | ~96.98 (Accuracy/Main Score) | Highly accurate binary classification |
| AIR-Bench (Zero-Shot) | 52.93 (Overall) | Outperforms many 7B embedders without examples |
| AIR-Bench (Few-Shot) | 54.36 (Overall) | State-of-the-Art performance; clear improvement via ICL |
4. Direct Comparative Analysis
| Feature / Attribute | bge-en-icl (BAAI) | e5-mistral-7b-instruct | bge-large-en-v1.5 |
|---|---|---|---|
| Base Architecture | Mistral-7B | Mistral-7B | BERT-Large |
| Parameter Count | ~7B | ~7B | ~335M |
| Embedding Size | 4096 dim | 4096 dim | 1024 dim |
| In-Context Learning | Yes (Dynamic few-shot) | Instruction-tuned (No ICL) | No |
| Compute Overhead | Very High (LLM-based) | Very High (LLM-based) | Low (Encoder-based) |
5. Hardware Requirements & Quantization Specs
| Deployment Format | File Size (approx.) | Minimum VRAM / RAM | Recommended Target Device |
|---|---|---|---|
| FP32 / FP16 (HuggingFace) | ~28.5 GB (FP32) | ~16 GB VRAM | 1x 24GB GPU (e.g., RTX 3090/4090) |
| FP16 Inference | ~14-15 GB | ~16 GB VRAM | High-end Consumer GPU / Server |
| Note: Due to being based on a 7B LLM, it requires significantly more compute and VRAM than traditional ~100M parameter encoder models. |
6. Recommended Inference Parameters & Usage
- Framework: Best utilized via
FlagEmbeddinglibrary (FlagICLModel) or HuggingFace Transformers. - Instruction Format: Requires task-specific query instructions (e.g.,
"Given a web search query, retrieve relevant passages that answer the query."). - ICL Examples Formulation: Provide examples as JSON-like dicts containing
instruct,query, andresponse. - FP16 Execution: Recommended (
use_fp16=True) to speed up computation with negligible performance degradation.
7. Core Strengths & Deployment Recommendations
- Key Strengths: Unprecedented adaptability to niche, domain-specific tasks using only a few examples in the prompt; exceptional baseline performance even in zero-shot scenarios.
- Known Limitations: Substantially higher latency and infrastructure costs compared to traditional BERT-based embedders (e.g.,
bge-base-en-v1.5); produces large 4096-dimensional vectors which increase vector database storage costs. - Best Use Cases: High-end RAG systems requiring extreme accuracy, complex classification or clustering tasks, and scenarios where the embedding task frequently changes or is highly specialized without resources for fine-tuning.