ModelsXiaomiMiMoMiMo-V2.5-tts-voiceclone
providerXiaomiMiMo /

MiMo-V2.5-tts-voiceclone

10 DZD/ 1m characters

Automatically convert input text into natural and fluent speech output. You can generate natural and vivid speech content by configuring parameters such as speech. Precisely replicate voices from audio samples to enable speech synthesis of any voice. style and voice.

Public
MiMo-V2.5-tts-voiceclone
ArchitectureDense
Context Windowspeech

MiMo-V2.5-TTS Series

Speech Synthesis (Text-to-Speech) supports automatically converting input text into natural and fluent speech output. You can generate natural and vivid speech content by configuring parameters such as speech style and voice.

Core Capabilities

  • Out-of-the-box built-in voices: A variety of high-quality built-in voices are available for quick use without additional configuration.
  • Voice design and cloning: Supports voice design via text description, or replication of arbitrary voices based on audio samples.
  • Diverse speech styles: Supports control over speed, emotion, role-play, dialects and other styles, for more vivid and natural speech expression.

List of Supported Models

Currently, three models of the MiMo-V2.5-TTS series are supported:

Model NameFunctionVoicePrecautions
MiMo-V2.5-TTSUse built-in high-quality voices for speech synthesisUse the high-quality voices from the built-in voices listSupports singing mode; does not support voice design and voice cloning
MiMo-V2.5-TTS-VoiceDesignCustomize voice through text descriptionAutomatically generate voices from text descriptions, without requiring presets or audio samplesDoes not support singing mode, built-in voices, or voice cloning
MiMo-V2.5-TTS-VoiceCloneReplicate any voice from audio samplesPrecisely replicate voices from audio samples to enable speech synthesis of any voiceDoes not support singing mode, built-in voices, or voice design

Style Control

The instruction-following ability of the model is sufficient to cover complex controls (a single natural language instruction is sufficient):

  • Multi-style Switching: A single character completes the style transition from announcement → whisper → roar within the same voice segment naturally.
  • Multi-emotion Mixing: Supports complex emotions such as "repressed anger", "smile with a sob", "gentle but tired", etc.
  • Multi-granularity control: From paragraph level (overall tone) → sentence level (rhythm) → word level (stress) → character granularity (choking, dragging, or breathy sound of a specific character).

Control Methods Placement

We currently offer two control methods:

  • Natural Language Control → Placed in role: user's content
  • Audio Tag Control → Placed in role: assistant's content

1. Natural Language Control

Enable the model to understand and generate speech in the corresponding style through natural language description. You can directly describe the desired speech style in a single sentence.

Example: Report good news to the leader in a brisk and upbeat tone, speaking at a slightly faster pace, with the uncontrollable excitement and a touch of pride after learning the results, and a bright and energetic voice.

Director Mode

For a more complex and refined control, use Director Mode to comprehensively depict characters and voices from three dimensions:

  • [Character] Clearly describe the character's identity, personality traits, physical appearance, and speaking habits.
  • [Scene] Describe what is happening at this moment, who you are talking to, and your emotional state (time, location, event, reactions).
  • [Guidance] Similar to acting instructions: speaking speed, breath control, pauses, accents, resonance position, timbre texture, and emotional fluctuations.

Director Mode Example: Role: The current head of the century-old noble Cen family. Since birth, she was adopted and raised by the gatekeeper of the ancestral temple, molded into a flawless, emotionless family totem. She has long lived in seclusion and has a strong sense of class alienation towards others. Scene: In the shadows of the ancestral hall, she watches the man who has broken through the security cordon at all costs to find her and attempts to elope with her. She will use the coldest and most rigid class barriers to strangle both the other person and the feelings that have just sprouted but are enough to start a prairie fire within herself. Guidance: A cold, languid yet extremely imposing deep-voiced mature woman... Extremely slow, with each word rolling on the tip of her tongue... A heavy and hard full voice, like a calm yet cold undercurrent...


2. Audio Tag Control

By embedding style tags and audio tags in the text, fine-grained control over speech can be directly achieved.

  • The overall style tag comes at the beginning (Style).
  • Fine-grained control tags can be inserted in the middle [Tag].
  • Supported bracket formats: Half-width (), full-width (), or [] can be used.
  • Format Example: (Style 1 Style 2)Content to be Synthesized

Style Tags

Style TypeStyle Example
Basic EmotionsHappy / Sad / Angry / Fearful / Amazed / Excited / Wronged / Calm / Indifferent
Complex EmotionsMelancholy / Relieved / Helpless / Guilty / Relieved / Jealous / Tired / Apprehensive / Emotional
Overall toneGentle / Cold / Lively / Serious / Lazy / Playful / Deep / Capable / Sharp
Timbre PositioningMagnetic / Mellow / Clear / Ethereal / Innocent / Old / Sweet / Hoarse / Elegant
Character ToneClamp voice / Big Sister voice / Shota voice / Uncle voice / Taiwanese accent
DialectNortheast dialect / Sichuan dialect / Henan dialect / Cantonese
Role-playingSun Wukong / Lin Daiyu
Singingsinging

Singing Mode Precaution: To experience a better singing style, you must add the (唱歌) or (sing) or (singing) tag at the very beginning of the target text. Lyrics are recommended to be in Chinese.

Style Tag Examples:

  • (Lazy) Let me sleep for five more minutes... just five minutes, really, for the last time.
  • (Cantonese) This is really amazing! Once you've tasted it, you won't forget!

Fine-Grained Audio Tags (Inserted in Text)

Style TypeStyle Example
Speech Rate & RhythmInhale / Take a deep breath / Sigh / Let out a long sigh / Pant / Hold one's breath
Emotional Statenervous / scared / excited / tired / wronged / coquettish / guilty / shocked / impatient
Speech FeaturesTrembling / Voice trembling / Pitch change / Cracked voice / Nasal voice / Breathiness / Hoarseness
Laughing & Crying ToneSmile / Chuckle / Laugh out loud / Sneer / Sob / Whimper / Choke / Wail

Fine-Grained Examples:

  • (nervously, takes a deep breath) Hoo... Calm down, calm down. It's just an interview... (speaking faster, muttering) I've rehearsed my self-introduction fifty times...
  • (Rapid breathing due to the cold) Hoo—hoo—This, this snow in the Greater Khingan Mountains... (cough) It can literally freeze one's bones...

Speech Synthesis Using Voice Cloning

By passing in audio samples, you can accurately replicate the target timbre and generate speech.

  • Supported Model: Currently, only the mimo-v2.5-tts-voiceclone model is supported.
  • Control: Supports controlling the style of synthesized speech by passing natural language instructions in the user message, or through audio tags.