all-MiniLM-L12-v2
3 DZD in / 1M tokens
all-MiniLM-L12-v2 is a widely used, highly efficient sentence embedding model developed by Sentence Transformers. Based on the 33.4-million parameter microsoft/MiniLM-L12-H384-uncased architecture, it maps sentences and short paragraphs into a dense 384-dimensional vector space. Fine-tuned on over 1 billion sentence pairs using self-supervised contrastive learning, it offers an excellent balance between retrieval speed, memory footprint, and semantic accuracy, making it a foundational baseline for lightweight semantic search and clustering applications.
Public

ArchitectureTransformer
Context Window512
1. Architectural & Technical Specifications
- Developer / Organization: Sentence Transformers (Hugging Face)
- Base Model & Lineage:
microsoft/MiniLM-L12-H384-uncased - Release Date: August 2021
- Model Architecture: Encoder-only Transformer (Bidirectional)
- Layer & Head Count: 12 Layers, 12 Attention heads
- Total Parameters: ~33.4 Million
- Active Parameters: ~33.4 Million
- Embedding Dimension: 384 dimensions
- Vocabulary Size: ~30,522 tokens
- Native Context Window: 256 tokens (hard truncation default)
- Max Output Length: Fixed-length 384-dimensional dense vector
- License: Apache 2.0 (Fully Permissive Open Source)
2. Training Data & Alignment Pipeline
- Dataset Composition: Fine-tuned on a massive dataset of 1 billion+ sentence pairs gathered from diverse sources including Reddit comments, Quora, StackExchange, Wikipedia, MS-MARCO, and SNLI/MNLI.
- Fine-Tuning Methodology: Self-supervised contrastive learning. The model was trained to predict which randomly sampled sentence was paired with the input sentence in the dataset.
- System Prompt / Prefix Conditioning: None required. The model natively processes raw text strings without task-specific prefixes or instructions.
- Similarity Metric: Cosine Similarity / Dot Product (when embeddings are normalized).
3. Benchmark Performance & Statistics
| Benchmark / Metric | Score / Result | Relative Performance vs Base/Competitors |
|---|---|---|
| MTEB Overall (English) | ~58.8 | Serves as the standard baseline for lightweight models |
| Semantic Search (STS) | ~68.7 | High baseline accuracy for short sentence similarity |
| Throughput / Latency | ~7500 sentences/sec | Extremely fast on both CPU and GPU hardware |
| Storage Footprint | ~120 MB | Minimal memory overhead, ideal for edge deployment |
4. Direct Comparative Analysis
| Feature / Attribute | all-MiniLM-L12-v2 | all-MiniLM-L6-v2 | e5-small-v2 | bge-small-en-v1.5 |
|---|---|---|---|---|
| Parameter Count | ~33.4M | ~22.7M | ~33.3M | ~33.4M |
| Embedding Size | 384 dim | 384 dim | 384 dim | 384 dim |
| Context Window | 256 tokens | 256 tokens | 512 tokens | 512 tokens |
| Layers | 12 | 6 | 12 | 12 |
| Prefix Required | No | No | Yes (query: / passage:) | Yes (Queries only) |
5. Hardware Requirements & Quantization Specs
| Deployment Format | File Size (approx.) | Minimum VRAM / RAM | Recommended Target Device |
|---|---|---|---|
| FP32 (PyTorch/HF) | ~134 MB | ~256 MB RAM / VRAM | CPU (x86_64, ARM, Raspberry Pi) |
| FP16 (GPU Optimized) | ~67 MB | ~128 MB VRAM | Any entry-level GPU (e.g., GTX 1060) |
| ONNX Runtime (FP32) | ~134 MB | ~256 MB RAM | Serverless Functions / Mobile Apps |
| INT8 / Quantized (ONNX) | ~34 MB | ~64 MB RAM | Browser (Transformers.js) / Edge IoT |
6. Recommended Inference Parameters & Usage
- Pooling Method: Mean pooling (Average pooling of all token embeddings, taking the attention mask into account).
- Normalization:
normalize_embeddings = True(Recommended for vector databases utilizing cosine similarity). - Input Formatting: Feed raw strings directly. No prefixes needed.
- Context Handling: By default, text longer than 256 tokens is silently truncated. Do not use this model for long documents without aggressive paragraph/sentence-level chunking.
- Framework Support: Natively integrated into
sentence-transformers, Hugging Facetransformers, andtransformers.jsfor in-browser execution.
7. Core Strengths & Deployment Recommendations
- Key Strengths: Ultra-lightweight footprint, blazing fast inference speeds even on CPU, completely open Apache 2.0 license, and exceptional integration ecosystem.
- Known Limitations: The strict 256-token limit makes it unsuitable for embedding full paragraphs or documents; its 384-dimensional output and older architecture limit absolute semantic accuracy compared to modern 1024-dimensional models (like
bge-large-en-v1.5). - Best Use Cases: Low-latency search systems, client-side browser embeddings (via WebAssembly), mobile device on-device retrieval, basic chat history similarity, and environments with strict compute or memory constraints.