MiMo-V2.5-tts-voiceclone
Automatically convert input text into natural and fluent speech output. You can generate natural and vivid speech content by configuring parameters such as speech. Precisely replicate voices from audio samples to enable speech synthesis of any voice. style and voice.

MiMo-V2.5-TTS Series
Speech Synthesis (Text-to-Speech) supports automatically converting input text into natural and fluent speech output. You can generate natural and vivid speech content by configuring parameters such as speech style and voice.
Core Capabilities
- Out-of-the-box built-in voices: A variety of high-quality built-in voices are available for quick use without additional configuration.
- Voice design and cloning: Supports voice design via text description, or replication of arbitrary voices based on audio samples.
- Diverse speech styles: Supports control over speed, emotion, role-play, dialects and other styles, for more vivid and natural speech expression.
List of Supported Models
Currently, three models of the MiMo-V2.5-TTS series are supported:
| Model Name | Function | Voice | Precautions |
|---|---|---|---|
| MiMo-V2.5-TTS | Use built-in high-quality voices for speech synthesis | Use the high-quality voices from the built-in voices list | Supports singing mode; does not support voice design and voice cloning |
| MiMo-V2.5-TTS-VoiceDesign | Customize voice through text description | Automatically generate voices from text descriptions, without requiring presets or audio samples | Does not support singing mode, built-in voices, or voice cloning |
| MiMo-V2.5-TTS-VoiceClone | Replicate any voice from audio samples | Precisely replicate voices from audio samples to enable speech synthesis of any voice | Does not support singing mode, built-in voices, or voice design |
Style Control
The instruction-following ability of the model is sufficient to cover complex controls (a single natural language instruction is sufficient):
- Multi-style Switching: A single character completes the style transition from announcement → whisper → roar within the same voice segment naturally.
- Multi-emotion Mixing: Supports complex emotions such as "repressed anger", "smile with a sob", "gentle but tired", etc.
- Multi-granularity control: From paragraph level (overall tone) → sentence level (rhythm) → word level (stress) → character granularity (choking, dragging, or breathy sound of a specific character).
Control Methods Placement
We currently offer two control methods:
- Natural Language Control → Placed in
role: user's content - Audio Tag Control → Placed in
role: assistant's content
1. Natural Language Control
Enable the model to understand and generate speech in the corresponding style through natural language description. You can directly describe the desired speech style in a single sentence.
Example: Report good news to the leader in a brisk and upbeat tone, speaking at a slightly faster pace, with the uncontrollable excitement and a touch of pride after learning the results, and a bright and energetic voice.
Director Mode
For a more complex and refined control, use Director Mode to comprehensively depict characters and voices from three dimensions:
- [Character] Clearly describe the character's identity, personality traits, physical appearance, and speaking habits.
- [Scene] Describe what is happening at this moment, who you are talking to, and your emotional state (time, location, event, reactions).
- [Guidance] Similar to acting instructions: speaking speed, breath control, pauses, accents, resonance position, timbre texture, and emotional fluctuations.
Director Mode Example: Role: The current head of the century-old noble Cen family. Since birth, she was adopted and raised by the gatekeeper of the ancestral temple, molded into a flawless, emotionless family totem. She has long lived in seclusion and has a strong sense of class alienation towards others. Scene: In the shadows of the ancestral hall, she watches the man who has broken through the security cordon at all costs to find her and attempts to elope with her. She will use the coldest and most rigid class barriers to strangle both the other person and the feelings that have just sprouted but are enough to start a prairie fire within herself. Guidance: A cold, languid yet extremely imposing deep-voiced mature woman... Extremely slow, with each word rolling on the tip of her tongue... A heavy and hard full voice, like a calm yet cold undercurrent...
2. Audio Tag Control
By embedding style tags and audio tags in the text, fine-grained control over speech can be directly achieved.
- The overall style tag comes at the beginning
(Style). - Fine-grained control tags can be inserted in the middle
[Tag]. - Supported bracket formats: Half-width
(), full-width(), or[]can be used. - Format Example:
(Style 1 Style 2)Content to be Synthesized
Style Tags
| Style Type | Style Example |
|---|---|
| Basic Emotions | Happy / Sad / Angry / Fearful / Amazed / Excited / Wronged / Calm / Indifferent |
| Complex Emotions | Melancholy / Relieved / Helpless / Guilty / Relieved / Jealous / Tired / Apprehensive / Emotional |
| Overall tone | Gentle / Cold / Lively / Serious / Lazy / Playful / Deep / Capable / Sharp |
| Timbre Positioning | Magnetic / Mellow / Clear / Ethereal / Innocent / Old / Sweet / Hoarse / Elegant |
| Character Tone | Clamp voice / Big Sister voice / Shota voice / Uncle voice / Taiwanese accent |
| Dialect | Northeast dialect / Sichuan dialect / Henan dialect / Cantonese |
| Role-playing | Sun Wukong / Lin Daiyu |
| Singing | singing |
Singing Mode Precaution: To experience a better singing style, you must add the
(唱歌)or(sing)or(singing)tag at the very beginning of the target text. Lyrics are recommended to be in Chinese.
Style Tag Examples:
(Lazy)Let me sleep for five more minutes... just five minutes, really, for the last time.(Cantonese)This is really amazing! Once you've tasted it, you won't forget!
Fine-Grained Audio Tags (Inserted in Text)
| Style Type | Style Example |
|---|---|
| Speech Rate & Rhythm | Inhale / Take a deep breath / Sigh / Let out a long sigh / Pant / Hold one's breath |
| Emotional State | nervous / scared / excited / tired / wronged / coquettish / guilty / shocked / impatient |
| Speech Features | Trembling / Voice trembling / Pitch change / Cracked voice / Nasal voice / Breathiness / Hoarseness |
| Laughing & Crying Tone | Smile / Chuckle / Laugh out loud / Sneer / Sob / Whimper / Choke / Wail |
Fine-Grained Examples:
(nervously, takes a deep breath)Hoo... Calm down, calm down. It's just an interview...(speaking faster, muttering)I've rehearsed my self-introduction fifty times...(Rapid breathing due to the cold)Hoo—hoo—This, this snow in the Greater Khingan Mountains...(cough)It can literally freeze one's bones...
Speech Synthesis Using Voice Cloning
By passing in audio samples, you can accurately replicate the target timbre and generate speech.
- Supported Model: Currently, only the
mimo-v2.5-tts-voiceclonemodel is supported. - Control: Supports controlling the style of synthesized speech by passing natural language instructions in the user message, or through audio tags.