providersentence-transformers /

all-MiniLM-L12-v2

3 DZD in / 1M tokens

all-MiniLM-L12-v2 is a widely used, highly efficient sentence embedding model developed by Sentence Transformers. Based on the 33.4-million parameter microsoft/MiniLM-L12-H384-uncased architecture, it maps sentences and short paragraphs into a dense 384-dimensional vector space. Fine-tuned on over 1 billion sentence pairs using self-supervised contrastive learning, it offers an excellent balance between retrieval speed, memory footprint, and semantic accuracy, making it a foundational baseline for lightweight semantic search and clustering applications.

Public
all-MiniLM-L12-v2
ArchitectureTransformer
Context Window512

1. Architectural & Technical Specifications

  • Developer / Organization: Sentence Transformers (Hugging Face)
  • Base Model & Lineage: microsoft/MiniLM-L12-H384-uncased
  • Release Date: August 2021
  • Model Architecture: Encoder-only Transformer (Bidirectional)
  • Layer & Head Count: 12 Layers, 12 Attention heads
  • Total Parameters: ~33.4 Million
  • Active Parameters: ~33.4 Million
  • Embedding Dimension: 384 dimensions
  • Vocabulary Size: ~30,522 tokens
  • Native Context Window: 256 tokens (hard truncation default)
  • Max Output Length: Fixed-length 384-dimensional dense vector
  • License: Apache 2.0 (Fully Permissive Open Source)

2. Training Data & Alignment Pipeline

  • Dataset Composition: Fine-tuned on a massive dataset of 1 billion+ sentence pairs gathered from diverse sources including Reddit comments, Quora, StackExchange, Wikipedia, MS-MARCO, and SNLI/MNLI.
  • Fine-Tuning Methodology: Self-supervised contrastive learning. The model was trained to predict which randomly sampled sentence was paired with the input sentence in the dataset.
  • System Prompt / Prefix Conditioning: None required. The model natively processes raw text strings without task-specific prefixes or instructions.
  • Similarity Metric: Cosine Similarity / Dot Product (when embeddings are normalized).

3. Benchmark Performance & Statistics

Benchmark / MetricScore / ResultRelative Performance vs Base/Competitors
MTEB Overall (English)~58.8Serves as the standard baseline for lightweight models
Semantic Search (STS)~68.7High baseline accuracy for short sentence similarity
Throughput / Latency~7500 sentences/secExtremely fast on both CPU and GPU hardware
Storage Footprint~120 MBMinimal memory overhead, ideal for edge deployment

4. Direct Comparative Analysis

Feature / Attributeall-MiniLM-L12-v2all-MiniLM-L6-v2e5-small-v2bge-small-en-v1.5
Parameter Count~33.4M~22.7M~33.3M~33.4M
Embedding Size384 dim384 dim384 dim384 dim
Context Window256 tokens256 tokens512 tokens512 tokens
Layers1261212
Prefix RequiredNoNoYes (query: / passage:)Yes (Queries only)

5. Hardware Requirements & Quantization Specs

Deployment FormatFile Size (approx.)Minimum VRAM / RAMRecommended Target Device
FP32 (PyTorch/HF)~134 MB~256 MB RAM / VRAMCPU (x86_64, ARM, Raspberry Pi)
FP16 (GPU Optimized)~67 MB~128 MB VRAMAny entry-level GPU (e.g., GTX 1060)
ONNX Runtime (FP32)~134 MB~256 MB RAMServerless Functions / Mobile Apps
INT8 / Quantized (ONNX)~34 MB~64 MB RAMBrowser (Transformers.js) / Edge IoT

6. Recommended Inference Parameters & Usage

  • Pooling Method: Mean pooling (Average pooling of all token embeddings, taking the attention mask into account).
  • Normalization: normalize_embeddings = True (Recommended for vector databases utilizing cosine similarity).
  • Input Formatting: Feed raw strings directly. No prefixes needed.
  • Context Handling: By default, text longer than 256 tokens is silently truncated. Do not use this model for long documents without aggressive paragraph/sentence-level chunking.
  • Framework Support: Natively integrated into sentence-transformers, Hugging Face transformers, and transformers.js for in-browser execution.

7. Core Strengths & Deployment Recommendations

  • Key Strengths: Ultra-lightweight footprint, blazing fast inference speeds even on CPU, completely open Apache 2.0 license, and exceptional integration ecosystem.
  • Known Limitations: The strict 256-token limit makes it unsuitable for embedding full paragraphs or documents; its 384-dimensional output and older architecture limit absolute semantic accuracy compared to modern 1024-dimensional models (like bge-large-en-v1.5).
  • Best Use Cases: Low-latency search systems, client-side browser embeddings (via WebAssembly), mobile device on-device retrieval, basic chat history similarity, and environments with strict compute or memory constraints.