all-MiniLM-L6-v2
3 DZD in / 1M tokens
all-MiniLM-L6-v2 is an ultra-lightweight, high-speed sentence embedding model developed by Sentence Transformers. Built on the 22.7-million parameter nreimers/MiniLM-L6-H384-uncased architecture (a 6-layer distilled version of BERT), it maps text into a compact 384-dimensional dense vector space. Trained on over 1 billion sentence pairs using self-supervised contrastive learning, it serves as the industry standard benchmark for edge deployments, in-browser inference, and real-time semantic search with minimal latency.
Public

ArchitectureTransformer
Context Window512
1. Architectural & Technical Specifications
- Developer / Organization: Sentence Transformers (Hugging Face / Nils Reimers)
- Base Model & Lineage:
nreimers/MiniLM-L6-H384-uncased(Distilled from BERT/RoBERTa) - Release Date: August 2021
- Model Architecture: Encoder-only Transformer (Bidirectional)
- Layer & Head Count: 6 Layers, 12 Attention heads
- Total Parameters: ~22.7 Million
- Active Parameters: ~22.7 Million
- Embedding Dimension: 384 dimensions
- Vocabulary Size: ~30,522 tokens (WordPiece)
- Native Context Window: 256 tokens (hard truncation default, max 512 with positional adjustments)
- Max Output Length: Fixed-length 384-dimensional dense vector
- License: Apache 2.0 (Permissive Open Source)
2. Training Data & Alignment Pipeline
- Dataset Composition: Fine-tuned on a massive corpus of over 1 billion sentence pairs from heterogeneous sources (Reddit comments, StackExchange, Yahoo Answers, MS-MARCO, Quora Question Pairs, SNLI/MNLI, and Wikipedia).
- Fine-Tuning Methodology: Self-supervised contrastive learning with cross-entropy loss over positive/negative paired samples; knowledge distillation from larger teacher models.
- System Prompt / Prefix Conditioning: None. Processes raw text directly without task prefixes or formatting wrappers.
- Similarity Metric: Cosine Similarity / Dot Product (when normalized).
3. Benchmark Performance & Statistics
| Benchmark / Metric | Score / Result | Relative Performance vs Base/Competitors |
|---|---|---|
| MTEB Overall (English) | ~56.26 | Strong baseline for sub-30M parameter class |
| MTEB Retrieval (BEIR) | ~41.95 | Fast baseline for lightweight lookup |
| MTEB Semantic Textual Similarity (STS) | ~78.90 | High correlation for short sentence pairs |
| Inference Throughput (CPU) | ~14,200 sentences/sec | ~2x faster than 12-layer variants |
4. Direct Comparative Analysis
| Feature / Attribute | all-MiniLM-L6-v2 | all-MiniLM-L12-v2 | bge-small-en-v1.5 | e5-small-v2 |
|---|---|---|---|---|
| Parameter Count | ~22.7M | ~33.4M | ~33.4M | ~33.3M |
| Layers | 6 | 12 | 12 | 12 |
| Embedding Size | 384 dim | 384 dim | 384 dim | 384 dim |
| Context Window | 256 tokens | 256 tokens | 512 tokens | 512 tokens |
| Inference Speed | Ultra Fast | Very Fast | Fast | Fast |
| Prefix Requirement | None | None | Query only | query: & passage: |
5. Hardware Requirements & Quantization Specs
| Deployment Format | File Size (approx.) | Minimum VRAM / RAM | Recommended Target Device |
|---|---|---|---|
| FP32 (PyTorch/HF) | ~91 MB | ~150 MB RAM / VRAM | CPU (x86_64, ARM, Raspberry Pi) |
| FP16 (GPU Optimized) | ~45 MB | ~100 MB VRAM | Entry GPU / Mobile SoC |
| ONNX Runtime (FP32/FP16) | ~45 - 91 MB | ~100 MB RAM | Serverless Functions / Cloudflare Workers |
| INT8 / Quantized (ONNX/Wasm) | ~23 MB | ~50 MB RAM | In-browser (Transformers.js) / IoT |
6. Recommended Inference Parameters & Usage
- Pooling Method: Mean pooling (Average pooling across non-masked token embeddings).
- Normalization:
normalize_embeddings = True(Standardizes vectors to unit length for inner product / cosine search). - Input Text: Supply raw strings directly without prefixes.
- Chunking Strategy: Chunk documents into 1–3 sentence segments (under 256 tokens) to avoid silent truncation.
- Framework Support: Natively supported in
sentence-transformers,transformers,fastembed, and client-side via@xenova/transformers(JavaScript).
7. Core Strengths & Deployment Recommendations
- Key Strengths: Negligible computational overhead, sub-millisecond latency on CPU, minimal vector database storage requirements (384 dimensions), and universal runtime compatibility.
- Known Limitations: 256-token context ceiling limits complex document retrieval; lower semantic depth and nuance compared to larger 768/1024-dim models (e.g.,
bge-baseorbge-large). - Best Use Cases: Real-time search autocomplete, duplicate detection, client-side vector search in browser extensions/mobile apps, local chat memory indexing, and ultra-high-throughput serverless microservices.