all-mpnet-base-v2
3 DZD in / 1M tokens
all-mpnet-base-v2 is an open-source sentence embedding model developed by Sentence Transformers, built upon Microsoft's 109-million parameter MPNet (Masked and Permuted Pre-training) architecture. By combining the strengths of masked language modeling (BERT) and permuted language modeling (XLNet), it maps text into a dense 768-dimensional vector space. Fine-tuned on over 1 billion sentence pairs using contrastive learning, it long served as the primary quality gold standard among Sentence Transformers for semantic similarity, classification, and retrieval tasks.
Public

ArchitectureTransformer
Context Window512
1. Architectural & Technical Specifications
- Developer / Organization: Sentence Transformers (Hugging Face / Microsoft Research)
- Base Model & Lineage:
microsoft/mpnet-base(MPNet Architecture) - Release Date: August 2021
- Model Architecture: Encoder-only Transformer combining Masked Language Modeling (MLM) and Permuted Language Modeling (PLM)
- Layer & Head Count: 12 Layers, 12 Attention heads
- Total Parameters: ~109 Million
- Active Parameters: ~109 Million
- Embedding Dimension: 768 dimensions
- Vocabulary Size: ~30,527 tokens (WordPiece)
- Native Context Window: 384 tokens (default truncation limit in
sentence-transformers, architectural max is 514 tokens) - Max Output Length: Fixed-length 768-dimensional dense vector
- License: Apache 2.0 (Permissive Open Source)
2. Training Data & Alignment Pipeline
- Dataset Composition: Fine-tuned on a massive corpus of 1 billion+ sentence pairs across diverse web sources (Reddit, StackExchange, Yahoo Answers, MS-MARCO, Quora Question Pairs, SNLI/MNLI, SQuAD, and Wikipedia).
- Fine-Tuning Methodology: Self-supervised contrastive learning with multiple negative ranking loss (MNRL), trained to distinguish true sentence pairings from randomly sampled negatives.
- System Prompt / Prefix Conditioning: None required. Operates directly on raw text strings without task-specific prefixes or instruction templates.
- Similarity Metric: Cosine Similarity / Dot Product (when normalized).
3. Benchmark Performance & Statistics
| Benchmark / Metric | Score / Result | Relative Performance vs Base/Competitors |
|---|---|---|
| MTEB Overall (English) | ~57.78 | Historic benchmark leader among 100M parameter encoders |
| MTEB Retrieval (BEIR) | ~43.81 | Solid baseline for dense text retrieval |
| MTEB Semantic Textual Similarity (STS) | ~80.28 | Exceptional semantic alignment for sentence pairs |
| Inference Throughput (CPU) | ~2,800 sentences/sec | Slower than MiniLM due to 768-dim hidden states, but higher accuracy |
4. Direct Comparative Analysis
| Feature / Attribute | all-mpnet-base-v2 | all-MiniLM-L6-v2 | bge-base-en-v1.5 | e5-base-v2 |
|---|---|---|---|---|
| Base Architecture | MPNet-base | MiniLM-L6 (BERT) | BERT-base | BERT-base |
| Parameter Count | ~109M | ~22.7M | ~109M | ~109M |
| Embedding Size | 768 dim | 384 dim | 768 dim | 768 dim |
| Context Window | 384 tokens | 256 tokens | 512 tokens | 512 tokens |
| MTEB Score | ~57.8 | ~56.3 | ~63.5 | ~61.5 |
| Prefix Required | None | None | Query only | query: & passage: |
5. Hardware Requirements & Quantization Specs
| Deployment Format | File Size (approx.) | Minimum VRAM / RAM | Recommended Target Device |
|---|---|---|---|
| FP32 (PyTorch/HF) | ~438 MB | ~1 GB RAM / VRAM | Standard CPU (x86_64, ARM) / Entry Server |
| FP16 (GPU Optimized) | ~219 MB | ~512 MB VRAM | NVIDIA T4 / RTX 3060 / A10G |
| ONNX Runtime (FP32/FP16) | ~220 - 438 MB | ~500 MB RAM | Production microservices / High-throughput APIs |
| INT8 / Quantized (ONNX/Q8) | ~110 MB | ~256 MB RAM | Serverless containers / Low-memory cloud nodes |
6. Recommended Inference Parameters & Usage
- Pooling Method: Mean pooling (Average pooling across all non-padded token representations).
- Normalization:
normalize_embeddings = True(Standardizes vectors to unit length for inner product / cosine similarity). - Input Text: Feed raw strings directly without prefixes.
- Chunking Strategy: Chunk documents into paragraphs of 200–350 tokens to avoid exceeding the 384-token boundary.
- Framework Support: Universal support across
sentence-transformers,transformers,onnxruntime, LangChain, and LlamaIndex.
7. Core Strengths & Deployment Recommendations
- Key Strengths: Highly consistent semantic representations, robust zero-shot generalization across classic NLP tasks, completely prefix-free deployment, and extensive library support.
- Known Limitations: Context window truncated at 384 tokens by default; superseded in dense retrieval and MTEB performance by modern embedding architectures (like
bge-base-en-v1.5andgte-base). - Best Use Cases: Semantic search in legacy and mature Python ecosystems, text clustering, deduplication pipelines, sentence-level intent classification, and standard document similarity matching.