providersentence-transformers /

all-mpnet-base-v2

3 DZD in / 1M tokens

all-mpnet-base-v2 is an open-source sentence embedding model developed by Sentence Transformers, built upon Microsoft's 109-million parameter MPNet (Masked and Permuted Pre-training) architecture. By combining the strengths of masked language modeling (BERT) and permuted language modeling (XLNet), it maps text into a dense 768-dimensional vector space. Fine-tuned on over 1 billion sentence pairs using contrastive learning, it long served as the primary quality gold standard among Sentence Transformers for semantic similarity, classification, and retrieval tasks.

Public
all-mpnet-base-v2
ArchitectureTransformer
Context Window512

1. Architectural & Technical Specifications

  • Developer / Organization: Sentence Transformers (Hugging Face / Microsoft Research)
  • Base Model & Lineage: microsoft/mpnet-base (MPNet Architecture)
  • Release Date: August 2021
  • Model Architecture: Encoder-only Transformer combining Masked Language Modeling (MLM) and Permuted Language Modeling (PLM)
  • Layer & Head Count: 12 Layers, 12 Attention heads
  • Total Parameters: ~109 Million
  • Active Parameters: ~109 Million
  • Embedding Dimension: 768 dimensions
  • Vocabulary Size: ~30,527 tokens (WordPiece)
  • Native Context Window: 384 tokens (default truncation limit in sentence-transformers, architectural max is 514 tokens)
  • Max Output Length: Fixed-length 768-dimensional dense vector
  • License: Apache 2.0 (Permissive Open Source)

2. Training Data & Alignment Pipeline

  • Dataset Composition: Fine-tuned on a massive corpus of 1 billion+ sentence pairs across diverse web sources (Reddit, StackExchange, Yahoo Answers, MS-MARCO, Quora Question Pairs, SNLI/MNLI, SQuAD, and Wikipedia).
  • Fine-Tuning Methodology: Self-supervised contrastive learning with multiple negative ranking loss (MNRL), trained to distinguish true sentence pairings from randomly sampled negatives.
  • System Prompt / Prefix Conditioning: None required. Operates directly on raw text strings without task-specific prefixes or instruction templates.
  • Similarity Metric: Cosine Similarity / Dot Product (when normalized).

3. Benchmark Performance & Statistics

Benchmark / MetricScore / ResultRelative Performance vs Base/Competitors
MTEB Overall (English)~57.78Historic benchmark leader among 100M parameter encoders
MTEB Retrieval (BEIR)~43.81Solid baseline for dense text retrieval
MTEB Semantic Textual Similarity (STS)~80.28Exceptional semantic alignment for sentence pairs
Inference Throughput (CPU)~2,800 sentences/secSlower than MiniLM due to 768-dim hidden states, but higher accuracy

4. Direct Comparative Analysis

Feature / Attributeall-mpnet-base-v2all-MiniLM-L6-v2bge-base-en-v1.5e5-base-v2
Base ArchitectureMPNet-baseMiniLM-L6 (BERT)BERT-baseBERT-base
Parameter Count~109M~22.7M~109M~109M
Embedding Size768 dim384 dim768 dim768 dim
Context Window384 tokens256 tokens512 tokens512 tokens
MTEB Score~57.8~56.3~63.5~61.5
Prefix RequiredNoneNoneQuery onlyquery: & passage:

5. Hardware Requirements & Quantization Specs

Deployment FormatFile Size (approx.)Minimum VRAM / RAMRecommended Target Device
FP32 (PyTorch/HF)~438 MB~1 GB RAM / VRAMStandard CPU (x86_64, ARM) / Entry Server
FP16 (GPU Optimized)~219 MB~512 MB VRAMNVIDIA T4 / RTX 3060 / A10G
ONNX Runtime (FP32/FP16)~220 - 438 MB~500 MB RAMProduction microservices / High-throughput APIs
INT8 / Quantized (ONNX/Q8)~110 MB~256 MB RAMServerless containers / Low-memory cloud nodes

6. Recommended Inference Parameters & Usage

  • Pooling Method: Mean pooling (Average pooling across all non-padded token representations).
  • Normalization: normalize_embeddings = True (Standardizes vectors to unit length for inner product / cosine similarity).
  • Input Text: Feed raw strings directly without prefixes.
  • Chunking Strategy: Chunk documents into paragraphs of 200–350 tokens to avoid exceeding the 384-token boundary.
  • Framework Support: Universal support across sentence-transformers, transformers, onnxruntime, LangChain, and LlamaIndex.

7. Core Strengths & Deployment Recommendations

  • Key Strengths: Highly consistent semantic representations, robust zero-shot generalization across classic NLP tasks, completely prefix-free deployment, and extensive library support.
  • Known Limitations: Context window truncated at 384 tokens by default; superseded in dense retrieval and MTEB performance by modern embedding architectures (like bge-base-en-v1.5 and gte-base).
  • Best Use Cases: Semantic search in legacy and mature Python ecosystems, text clustering, deduplication pipelines, sentence-level intent classification, and standard document similarity matching.