providersentence-transformers /

all-MiniLM-L6-v2

3 DZD in / 1M tokens

all-MiniLM-L6-v2 is an ultra-lightweight, high-speed sentence embedding model developed by Sentence Transformers. Built on the 22.7-million parameter nreimers/MiniLM-L6-H384-uncased architecture (a 6-layer distilled version of BERT), it maps text into a compact 384-dimensional dense vector space. Trained on over 1 billion sentence pairs using self-supervised contrastive learning, it serves as the industry standard benchmark for edge deployments, in-browser inference, and real-time semantic search with minimal latency.

Public
all-MiniLM-L6-v2
ArchitectureTransformer
Context Window512

1. Architectural & Technical Specifications

  • Developer / Organization: Sentence Transformers (Hugging Face / Nils Reimers)
  • Base Model & Lineage: nreimers/MiniLM-L6-H384-uncased (Distilled from BERT/RoBERTa)
  • Release Date: August 2021
  • Model Architecture: Encoder-only Transformer (Bidirectional)
  • Layer & Head Count: 6 Layers, 12 Attention heads
  • Total Parameters: ~22.7 Million
  • Active Parameters: ~22.7 Million
  • Embedding Dimension: 384 dimensions
  • Vocabulary Size: ~30,522 tokens (WordPiece)
  • Native Context Window: 256 tokens (hard truncation default, max 512 with positional adjustments)
  • Max Output Length: Fixed-length 384-dimensional dense vector
  • License: Apache 2.0 (Permissive Open Source)

2. Training Data & Alignment Pipeline

  • Dataset Composition: Fine-tuned on a massive corpus of over 1 billion sentence pairs from heterogeneous sources (Reddit comments, StackExchange, Yahoo Answers, MS-MARCO, Quora Question Pairs, SNLI/MNLI, and Wikipedia).
  • Fine-Tuning Methodology: Self-supervised contrastive learning with cross-entropy loss over positive/negative paired samples; knowledge distillation from larger teacher models.
  • System Prompt / Prefix Conditioning: None. Processes raw text directly without task prefixes or formatting wrappers.
  • Similarity Metric: Cosine Similarity / Dot Product (when normalized).

3. Benchmark Performance & Statistics

Benchmark / MetricScore / ResultRelative Performance vs Base/Competitors
MTEB Overall (English)~56.26Strong baseline for sub-30M parameter class
MTEB Retrieval (BEIR)~41.95Fast baseline for lightweight lookup
MTEB Semantic Textual Similarity (STS)~78.90High correlation for short sentence pairs
Inference Throughput (CPU)~14,200 sentences/sec~2x faster than 12-layer variants

4. Direct Comparative Analysis

Feature / Attributeall-MiniLM-L6-v2all-MiniLM-L12-v2bge-small-en-v1.5e5-small-v2
Parameter Count~22.7M~33.4M~33.4M~33.3M
Layers6121212
Embedding Size384 dim384 dim384 dim384 dim
Context Window256 tokens256 tokens512 tokens512 tokens
Inference SpeedUltra FastVery FastFastFast
Prefix RequirementNoneNoneQuery onlyquery: & passage:

5. Hardware Requirements & Quantization Specs

Deployment FormatFile Size (approx.)Minimum VRAM / RAMRecommended Target Device
FP32 (PyTorch/HF)~91 MB~150 MB RAM / VRAMCPU (x86_64, ARM, Raspberry Pi)
FP16 (GPU Optimized)~45 MB~100 MB VRAMEntry GPU / Mobile SoC
ONNX Runtime (FP32/FP16)~45 - 91 MB~100 MB RAMServerless Functions / Cloudflare Workers
INT8 / Quantized (ONNX/Wasm)~23 MB~50 MB RAMIn-browser (Transformers.js) / IoT

6. Recommended Inference Parameters & Usage

  • Pooling Method: Mean pooling (Average pooling across non-masked token embeddings).
  • Normalization: normalize_embeddings = True (Standardizes vectors to unit length for inner product / cosine search).
  • Input Text: Supply raw strings directly without prefixes.
  • Chunking Strategy: Chunk documents into 1–3 sentence segments (under 256 tokens) to avoid silent truncation.
  • Framework Support: Natively supported in sentence-transformers, transformers, fastembed, and client-side via @xenova/transformers (JavaScript).

7. Core Strengths & Deployment Recommendations

  • Key Strengths: Negligible computational overhead, sub-millisecond latency on CPU, minimal vector database storage requirements (384 dimensions), and universal runtime compatibility.
  • Known Limitations: 256-token context ceiling limits complex document retrieval; lower semantic depth and nuance compared to larger 768/1024-dim models (e.g., bge-base or bge-large).
  • Best Use Cases: Real-time search autocomplete, duplicate detection, client-side vector search in browser extensions/mobile apps, local chat memory indexing, and ultra-high-throughput serverless microservices.