bge-m3-multi
5 DZD in / 1M tokens
bge-m3-multi is an alternative designation for the bge-m3 embedding model developed by the Beijing Academy of Artificial Intelligence (BAAI). Built on a 567-million parameter XLM-RoBERTa architecture, the "Multi" suffix emphasizes its core design pillars: Multi-Linguality (supporting over 100 languages), Multi-Functionality (simultaneous Dense, Sparse, and Multi-Vector ColBERT outputs), and Multi-Granularity (handling inputs up to 512 tokens). It serves as a unified, state-of-the-art engine for hybrid search and complex multilingual Retrieval-Augmented Generation (RAG) pipelines.
Public

ArchitectureTransformer
Context Window512
1. Overview
- Model Name: bge-m3
- Developer: BAAI (Beijing Academy of Artificial Intelligence)
- Model Type: Advanced Text Embedding Model
- The "M3" Core Concept:
- Multi-Lingual: Native support for over 100 languages.
- Multi-Function: Supports Dense, Sparse (Lexical), and Multi-vector (ColBERT-style) retrieval.
- Multi-Granularity: Handles input lengths from short sentences up to 512-token passages.
2. Technical Specifications
- Architecture: Transformer-based (RetroMAE pre-training methodology)
- Max Context Length: Up to 512 tokens
- Embedding Dimension: 1024
- Parameters: ~567 Million
- Output Formats:
- Dense: Standard single vector representing the whole text.
- Sparse: Lexical weights (similar to BM25, good for keyword matching).
- Multi-Vector: Token-level embeddings (ColBERT-style for fine-grained alignment).
3. Performance & Capabilities
- Cross-Lingual Retrieval: Highly optimized for matching queries in one language (e.g., Arabic) with documents in another (e.g., English).
- Long-Document Processing: Accepts up to 512 tokens per input; longer inputs are truncated, so split long documents into chunks before embedding.
4. Hardware & Integration
- Resource Requirements: Lightweight. Can run efficiently on CPUs for smaller batches, or entry-level GPUs (T4, L4) for high-throughput production.
- Framework Compatibility: HuggingFace
sentence-transformers, LlamaIndex, LangChain, FlagEmbedding. - Database Fit: Ideal for Vector DBs like PostgreSQL (Supabase
pgvector), Pinecone, or Qdrant.
5. Primary Use Cases
- High-accuracy RAG (Retrieval-Augmented Generation) pipelines.
- Hybrid Search implementations (combining dense semantic search with sparse keyword search).
- Multi-lingual enterprise search platforms.