Qwen Audio 3.1 TTS Flash Speech Synthesis
- Convert text to speech, up to
5000characters per request - Choose from 68 system voices; see the voice list. You can also use a voice you created with Voice Enrollment, through cloning or design
- Custom voices expire by default
6 hoursafter the creation task completes; synthesis after expiry returns404(voice_expired). Create a new voice with Voice Enrollment - Supports natural-language instructions (
instruction), inline emotion and paralinguistic tags, SSML, and custom pronunciation (hot_fix) - Only the parameters listed below are accepted; other parameters return
400(unsupported_parameter) - Asynchronous processing; use the returned task ID to query the result
- Tasks cannot be canceled after submission
- Generated audio links are valid for 24 hours; save them promptly
Billing:
- Billed by actual token usage: input tokens (related to text length) and output tokens (related to generated audio duration) are priced separately
- Credits are reserved at submission based on text length (
usage.credits_reserved). On completion, billing is settled against actual usage: excess is refunded and any shortfall is charged. Failed tasks receive a full refund - Numbers, letters, and symbols may be read individually, producing much longer audio and more output tokens than ordinary prose of the same length; actual usage may exceed the reservation
- For the same text, instructions or tags that slow delivery (such as
[very slowly]) can substantially increase output tokens. Adjustingspeech_rateand inserting SSML pauses do not increase output tokens
Task result (when status is completed):
| Field | Description |
|---|---|
results[0] | Audio URL |
result_data[0].audio_url | Audio URL, identical to results[0] |
result_data[0].format | Audio format |
result_data[0].sample_rate | Sample rate (Hz) |
usage.input_tokens / usage.output_tokens / usage.total_tokens | Token usage for this synthesis |
usage.credits_used | Actual credits used |
Authorizations
Bearer token authentication is required for all endpoints
Get an API key:
Visit API Key Management to obtain your API key
Add this request header:
Authorization: Bearer YOUR_API_KEY
Body
- Option 1
- Option 2
Provide at least one non-empty text field, prompt or input; if both are supplied, their contents must match. If both response_format and format are supplied, their values must match.
Text to synthesize
Constraints:
- Maximum
5000characters - Keep punctuation in long text: long continuous passages without sentence breaks may be truncated upstream at about
1500output tokens (about 120 seconds of audio). The task still reports success and is billed for the tokens actually generated; the gateway cannot detect this truncation - You can also use
input. At least one non-empty text field is required; supplying just one is recommended. Different contents in both fields return400(parameter_conflict) - Text must be in a language supported by the selected voice; otherwise pronunciation may be incorrect
Emotion and paralinguistic tags: Embed tags directly in the text without extra parameters; tag text counts toward billing characters
- Control tags: Set the emotion or style of the following text until the next control tag.
[sad]sad,[amazed]amazed,[deep and loud shouting]deep, loud shouting,[trembling]trembling,[angry]angry,[excited]excited,[sarcastic]sarcastic,[curious]curious,[like dracula]low and eerie,[bored]bored,[tired]tired,[scornful]scornful,[shouting]shouting,[asmr]soft ASMR whispers,[panicked]panicked,[mischievously]mischievous,[empathetic]empathetic,[whispers]whispering,[reluctantly]reluctant,[crying]crying,[serious]serious,[very slowly]very slowly,[very fast]very fast - Paralinguistic tags: Insert a vocal effect at this position without changing the emotion of the surrounding text.
[gasp]gasp,[sighing]sigh,[clears throat]clear the throat,[giggles]giggle,[laughing]laugh,[cough]cough,[snorts]snort
Example: [excited]今天的天气真不错![laughing]我们一起出去玩吧!
When enable_ssml is true, this field is parsed as SSML
5000\S"我家的后面有一个很大的花园。"
Model name
qwen-audio-3.1-tts-flash "qwen-audio-3.1-tts-flash"
Alias for prompt, with the same length limits and usage rules
- Provide at least one non-empty text field,
promptorinput - If both are supplied, their contents must match; otherwise
400(parameter_conflict) is returned
5000"我家的后面有一个很大的花园。"
Voice name, case-sensitive
- 68 system voices; see the voice list for names, gender, and use cases
- Defaults to
longanhuan_v3.1if omitted - You can also supply a voice you created with Voice Enrollment: cloned voices follow
qwen-audio-3.1-tts-flash-{prefix}-{32-character-id}, and designed voices followqwen-audio-3.1-tts-flash-vd-{prefix}-{32-character-id}. Only the account that created the voice can use it. Voices from other models, such asqwen-tts-vd-…fromqwen-voice-design, return400(invalid_voice). Missing voices or voices owned by another account return404(voice_not_found) - Custom voices expire by default
6 hoursafter the creation task completes; synthesis after expiry returns404(voice_expired). Create a new voice with Voice Enrollment
"longanhuan_v3.1"
Output audio format: mp3, wav, or opus; defaults to mp3
opususes an Ogg Opus container- You can also use
format. Supplying just one field is recommended; conflicting values return400(parameter_conflict)
mp3, wav, opus "mp3"
Alias for response_format; supports mp3, wav, and opus
- Defaults to
mp3when neither field is supplied - If both fields are supplied, their values must match; otherwise
400(parameter_conflict) is returned
mp3, wav, opus "mp3"
Output sample rate (Hz)
22050and44100are not supported whenresponse_formatisopus- Omitted or
nulluses the default;0or a value outside the list returns400
8000, 12000, 16000, 22050, 24000, 44100, 48000, null 24000
Volume, from 0 to 100
0 <= x <= 10050
Speech rate multiplier
1.0: normal speed (default)2.0: double speed;0.5: half speed
Range: 0.5 to 2.0. Adjusting speech rate does not change output token count
0.5 <= x <= 21
Pitch multiplier
1.0: default pitch- Above
1.0raises pitch; below1.0lowers it
Range: 0.5 to 2.0
Changing pitch also changes speech rate and audio duration
- Higher pitch speeds up speech and shortens audio; lower pitch slows it down and lengthens audio. Duration varies approximately inversely with the square of the pitch value
- For a sentence lasting about 2.8 seconds at
1.0:0.8is about 4.3 seconds,1.2about 2.1 seconds,0.5about 10.9 seconds, and2.0about 0.7 seconds - Small adjustments between
0.8and1.2are recommended; values near0.5or2.0make speech noticeably too slow or too fast - If
speech_rateis also supplied with a value other than1.0,pitchhas no effect; the two cannot be combined - Adjusting pitch does not change output token count
0.5 <= x <= 21
Natural-language instructions to control emotion, tone, character, dialect, and other delivery choices
Constraints:
- Maximum
100billing characters: Han characters (including Japanese kanji and Korean hanja) count as 2; other characters count as 1, including kana and Hangul (about 50 Han characters or 100 English characters). Exceeding the limit returns400
Examples:
用欢快、热情的语气说(speak cheerfully and enthusiastically)请用上海话表达(use Shanghainese; multilingual and dialect voices)Speak slowly in a calm and gentle tone
Instructions do not count toward input tokens, but can change the generated audio duration and therefore output tokens
The parameter is
instruction(singular);instructionsreturns400
"用欢快、热情的语气说"
Target-language hint to improve the reading of numbers, abbreviations, and symbols, and synthesis in less common languages
For example, with zh, 110 in hello, this is 110 is read as “yao yao ling” in Chinese
| Value | Language | Value | Language |
|---|---|---|---|
zh | Chinese | th | Thai |
en | English | id | Indonesian |
fr | French | vi | Vietnamese |
de | German | es | Spanish |
ja | Japanese | it | Italian |
ko | Korean | ms | Malay |
ru | Russian | fil | Filipino |
pt | Portuguese | ar | Arabic |
The model detects the language if omitted; this parameter does not translate text
zh, en, fr, de, ja, ko, ru, pt, th, id, vi, es, it, ms, fil, ar "zh"
Parse prompt as SSML
When enabled, SSML tags can be used, such as <break time="1s"/> for a pause:
<speak>欢迎收听今天的节目。<break time="1s"/>我们马上开始。</speak>
Pauses inserted with SSML do not count toward output tokens
false
Custom pronunciation and text replacement to correct polyphonic characters, proper names, and other pronunciations
pronunciation: annotate words with pinyin (separate syllables with spaces and mark tones with digits, such astian1 qi4) to override the default pronunciationreplace: replace specified words before synthesis. Synthesis and billing use the replaced text, which must also stay within5000characters; exceeding the limit returns400(prompt_too_long)
The two lists may contain at most 200 entries in total, counted as key-value pairs within the objects. Exceeding the limit returns 400 (invalid_parameter)
Provide at least one list. Each provided list must be a non-empty array of objects in {"word": "value"} form
Example:
{ "pronunciation": [{"天气": "tian1 qi4"}], "replace": [{"今天": "金天"}] }
Show child attributes
Show child attributes
Embed an invisible AIGC marker in generated audio (applies to wav / mp3 / opus)
false
HTTPS callback URL for the task result
Timing:
- Triggered when the task completes (
completed) or fails (failed); this model does not support cancellation - Sent after billing has been confirmed
Security requirements:
- HTTPS only
- Private-network IPs are prohibited (127.0.0.1, 10.x.x.x, 172.16–31.x.x, 192.168.x.x, etc.)
- URL length at most
2048characters
Delivery:
- Timeout:
10seconds - At most
3retries after failure, with delays of1/2/4seconds - The callback body has the same format as the task query response
- A 2xx status is success; other statuses trigger retries
"https://your-domain.com/webhooks/tts-completed"
Response
Speech synthesis task created successfully
Task creation timestamp
1790000000
Task ID
"task-unified-1790000000-abcd1234"
Model actually used
"qwen-audio-3.1-tts-flash"
Specific task object type
audio.generation.task Task progress percentage (0–100)
0 <= x <= 1000
Task status
pending, processing, completed, failed "pending"
Audio task details
Show child attributes
Show child attributes
Task output type
audio "audio"
Usage and billing information
Show child attributes
Show child attributes