Skip to main content
POST

Authorizations

Authorization
string
header
required

Bearer token authentication is required for all endpoints

Get an API key:

Visit API Key Management to obtain your API key

Add this request header:

Body

application/json

Provide at least one non-empty text field, prompt or input; if both are supplied, their contents must match. If both response_format and format are supplied, their values must match.

prompt
string
required

Text to synthesize

Constraints:

  • Maximum 5000 characters
  • Keep punctuation in long text: long continuous passages without sentence breaks may be truncated upstream at about 1500 output tokens (about 120 seconds of audio). The task still reports success and is billed for the tokens actually generated; the gateway cannot detect this truncation
  • You can also use input. At least one non-empty text field is required; supplying just one is recommended. Different contents in both fields return 400 (parameter_conflict)
  • Text must be in a language supported by the selected voice; otherwise pronunciation may be incorrect

Emotion and paralinguistic tags: Embed tags directly in the text without extra parameters; tag text counts toward billing characters

  • Control tags: Set the emotion or style of the following text until the next control tag. [sad] sad, [amazed] amazed, [deep and loud shouting] deep, loud shouting, [trembling] trembling, [angry] angry, [excited] excited, [sarcastic] sarcastic, [curious] curious, [like dracula] low and eerie, [bored] bored, [tired] tired, [scornful] scornful, [shouting] shouting, [asmr] soft ASMR whispers, [panicked] panicked, [mischievously] mischievous, [empathetic] empathetic, [whispers] whispering, [reluctantly] reluctant, [crying] crying, [serious] serious, [very slowly] very slowly, [very fast] very fast
  • Paralinguistic tags: Insert a vocal effect at this position without changing the emotion of the surrounding text. [gasp] gasp, [sighing] sigh, [clears throat] clear the throat, [giggles] giggle, [laughing] laugh, [cough] cough, [snorts] snort

Example: [excited]今天的天气真不错![laughing]我们一起出去玩吧!

When enable_ssml is true, this field is parsed as SSML

Maximum string length: 5000
Pattern: \S
Example:

"我家的后面有一个很大的花园。"

model
enum<string>
default:qwen-audio-3.1-tts-flash
required

Model name

Available options:
qwen-audio-3.1-tts-flash
Example:

"qwen-audio-3.1-tts-flash"

input
string

Alias for prompt, with the same length limits and usage rules

  • Provide at least one non-empty text field, prompt or input
  • If both are supplied, their contents must match; otherwise 400 (parameter_conflict) is returned
Maximum string length: 5000
Example:

"我家的后面有一个很大的花园。"

voice
string
default:longanhuan_v3.1

Voice name, case-sensitive

  • 68 system voices; see the voice list for names, gender, and use cases
  • Defaults to longanhuan_v3.1 if omitted
  • You can also supply a voice you created with Voice Enrollment: cloned voices follow qwen-audio-3.1-tts-flash-{prefix}-{32-character-id}, and designed voices follow qwen-audio-3.1-tts-flash-vd-{prefix}-{32-character-id}. Only the account that created the voice can use it. Voices from other models, such as qwen-tts-vd-… from qwen-voice-design, return 400 (invalid_voice). Missing voices or voices owned by another account return 404 (voice_not_found)
  • Custom voices expire by default 6 hours after the creation task completes; synthesis after expiry returns 404 (voice_expired). Create a new voice with Voice Enrollment
Example:

"longanhuan_v3.1"

response_format
enum<string>
default:mp3

Output audio format: mp3, wav, or opus; defaults to mp3

  • opus uses an Ogg Opus container
  • You can also use format. Supplying just one field is recommended; conflicting values return 400 (parameter_conflict)
Available options:
mp3,
wav,
opus
Example:

"mp3"

format
enum<string>

Alias for response_format; supports mp3, wav, and opus

  • Defaults to mp3 when neither field is supplied
  • If both fields are supplied, their values must match; otherwise 400 (parameter_conflict) is returned
Available options:
mp3,
wav,
opus
Example:

"mp3"

sample_rate
enum<integer> | null
default:24000

Output sample rate (Hz)

  • 22050 and 44100 are not supported when response_format is opus
  • Omitted or null uses the default; 0 or a value outside the list returns 400
Available options:
8000,
12000,
16000,
22050,
24000,
44100,
48000,
null
Example:

24000

volume
integer
default:50

Volume, from 0 to 100

Required range: 0 <= x <= 100
Example:

50

speech_rate
number
default:1

Speech rate multiplier

  • 1.0: normal speed (default)
  • 2.0: double speed; 0.5: half speed

Range: 0.5 to 2.0. Adjusting speech rate does not change output token count

Required range: 0.5 <= x <= 2
Example:

1

pitch
number
default:1

Pitch multiplier

  • 1.0: default pitch
  • Above 1.0 raises pitch; below 1.0 lowers it

Range: 0.5 to 2.0

Changing pitch also changes speech rate and audio duration

  • Higher pitch speeds up speech and shortens audio; lower pitch slows it down and lengthens audio. Duration varies approximately inversely with the square of the pitch value
  • For a sentence lasting about 2.8 seconds at 1.0: 0.8 is about 4.3 seconds, 1.2 about 2.1 seconds, 0.5 about 10.9 seconds, and 2.0 about 0.7 seconds
  • Small adjustments between 0.8 and 1.2 are recommended; values near 0.5 or 2.0 make speech noticeably too slow or too fast
  • If speech_rate is also supplied with a value other than 1.0, pitch has no effect; the two cannot be combined
  • Adjusting pitch does not change output token count
Required range: 0.5 <= x <= 2
Example:

1

instruction
string

Natural-language instructions to control emotion, tone, character, dialect, and other delivery choices

Constraints:

  • Maximum 100 billing characters: Han characters (including Japanese kanji and Korean hanja) count as 2; other characters count as 1, including kana and Hangul (about 50 Han characters or 100 English characters). Exceeding the limit returns 400

Examples:

  • 用欢快、热情的语气说 (speak cheerfully and enthusiastically)
  • 请用上海话表达 (use Shanghainese; multilingual and dialect voices)
  • Speak slowly in a calm and gentle tone

Instructions do not count toward input tokens, but can change the generated audio duration and therefore output tokens

The parameter is instruction (singular); instructions returns 400

Example:

"用欢快、热情的语气说"

language
enum<string>

Target-language hint to improve the reading of numbers, abbreviations, and symbols, and synthesis in less common languages

For example, with zh, 110 in hello, this is 110 is read as “yao yao ling” in Chinese

The model detects the language if omitted; this parameter does not translate text

Available options:
zh,
en,
fr,
de,
ja,
ko,
ru,
pt,
th,
id,
vi,
es,
it,
ms,
fil,
ar
Example:

"zh"

enable_ssml
boolean
default:false

Parse prompt as SSML

When enabled, SSML tags can be used, such as <break time="1s"/> for a pause: <speak>欢迎收听今天的节目。<break time="1s"/>我们马上开始。</speak>

Pauses inserted with SSML do not count toward output tokens

Example:

false

hot_fix
object

Custom pronunciation and text replacement to correct polyphonic characters, proper names, and other pronunciations

  • pronunciation: annotate words with pinyin (separate syllables with spaces and mark tones with digits, such as tian1 qi4) to override the default pronunciation
  • replace: replace specified words before synthesis. Synthesis and billing use the replaced text, which must also stay within 5000 characters; exceeding the limit returns 400 (prompt_too_long)

The two lists may contain at most 200 entries in total, counted as key-value pairs within the objects. Exceeding the limit returns 400 (invalid_parameter)

Provide at least one list. Each provided list must be a non-empty array of objects in {"word": "value"} form

Example:

enable_aigc_tag
boolean
default:false

Embed an invisible AIGC marker in generated audio (applies to wav / mp3 / opus)

Example:

false

callback_url
string<uri>

HTTPS callback URL for the task result

Timing:

  • Triggered when the task completes (completed) or fails (failed); this model does not support cancellation
  • Sent after billing has been confirmed

Security requirements:

  • HTTPS only
  • Private-network IPs are prohibited (127.0.0.1, 10.x.x.x, 172.16–31.x.x, 192.168.x.x, etc.)
  • URL length at most 2048 characters

Delivery:

  • Timeout: 10 seconds
  • At most 3 retries after failure, with delays of 1 / 2 / 4 seconds
  • The callback body has the same format as the task query response
  • A 2xx status is success; other statuses trigger retries
Example:

"https://your-domain.com/webhooks/tts-completed"

Response

Speech synthesis task created successfully

created
integer

Task creation timestamp

Example:

1790000000

id
string

Task ID

Example:

"task-unified-1790000000-abcd1234"

model
string

Model actually used

Example:

"qwen-audio-3.1-tts-flash"

object
enum<string>

Specific task object type

Available options:
audio.generation.task
progress
integer

Task progress percentage (0–100)

Required range: 0 <= x <= 100
Example:

0

status
enum<string>

Task status

Available options:
pending,
processing,
completed,
failed
Example:

"pending"

task_info
object

Audio task details

type
enum<string>

Task output type

Available options:
audio
Example:

"audio"

usage
object

Usage and billing information