Skip to main content
Qwen Audio 3.1 TTS Flash selects a voice with the voice parameter in the speech synthesis API. There are 68 system voices, ready to use without creating a voice first.
  • voice is case-sensitive; copy its value exactly as shown below
  • If voice is omitted, the default is longanhuan_v3.1
  • Use the table’s voice column in requests; names, gender, voice qualities, and use cases help you choose a voice
  • The text must be in a language supported by the selected voice; otherwise pronunciation may be incorrect or speech unnatural

Usage Examples

Send the following request body to POST /v1/audios/generations, replacing voice with the desired voice ID. See the speech synthesis API for all parameters and task queries.
All voices can use instruction and inline emotion tags, such as [excited] and [laughing], to control emotion and tone. See the prompt and instruction descriptions in the speech synthesis API.

Multilingual and Dialect Voices

4 voices. Each supports all the dialects and languages below:
  • Dialects: Shanghainese, Cantonese, Northeastern Chinese, Chongqing, Shaanxi, Yunnan, Ningbo, and Gansu dialects
  • Languages: Japanese, Korean, French, German, Portuguese, Italian, Vietnamese, and Indonesian
For dialect synthesis, specify the dialect in instruction, for example 请用上海话表达 (use Shanghainese).

Premium Chinese Voices

22 voices, supporting Mandarin Chinese only.

Premium English Voices

15 voices, supporting English only.

Other System Voices

27 voices. Supported languages are not listed individually for this group; use text in your target language to check the synthesis quality.

Custom Voices

For your own voice, use Voice Enrollment: provide a recording to clone a voice, or a text description to design one. After the task completes, pass the full returned voice name to the speech synthesis API.
  • Custom voices can only be used by the account that created them
  • By default, they expire 6 hours after the creation task completes; synthesis calls do not extend their lifetime
  • After expiry, synthesis returns 404 (voice_expired); create a new voice with Voice Enrollment
  • Voices created with the old qwen-voice-design model cannot be used here; recreate them with Voice Enrollment when migrating