Skip to main content
POST

Authorizations

Authorization
string
header
required

Bearer token authentication is required for all endpoints

Get an API key:

Visit API Key Management to obtain your API key

Add this request header:

Body

application/json

Create a voice from a recording of a person. Providing audio_url selects voice cloning; omit the design fields voice_prompt, preview_text, sample_rate, and response_format. Non-empty design text, a non-zero sample rate, or a non-empty format returns 400; sample_rate: 0/null and response_format: ""/null are treated as omitted.

model
enum<string>
default:voice-enrollment
required

Model name

Available options:
voice-enrollment
Example:

"voice-enrollment"

audio_url
string<uri>
required

URL of the recording to clone. This field selects voice cloning and cannot be supplied together with voice_prompt

URL requirements:

  • HTTP or HTTPS, publicly accessible without authentication
  • Local or private-network addresses return 400 (invalid_media_url)
  • Maximum 2048 characters
  • Non-HTTP(S) protocols such as FTP return 400 (invalid_media_url)
  • URLs only, not Base64; a Base64 data URI returns 400 (invalid_parameter)

Audio requirements (the task may fail if unmet):

  • WAV, MP3, or M4A
  • At most 60 seconds; insufficient voiced content can also fail (a 2-second recording failed in testing)
  • File size at most 10 MB
  • Sample rate of at least 16 kHz
  • Must contain clear human speech; silence, music, and other non-speech audio are rejected

Recording recommendations:

  • 10–20 seconds, including at least 5 seconds of continuous, clear reading, with pauses no longer than 2 seconds
  • Mono; for stereo recordings, only the first channel is used
  • No background music, noise, or other voices; speak normally, do not sing

If audio cannot be downloaded or does not meet the requirements, the task fails and credits are fully refunded

Only clone your own voice or a voice you have explicit permission to clone

Maximum string length: 2048
Pattern: [^\u0009-\u000D\u0020\u0085\u00A0\u1680\u2000-\u200A\u2028\u2029\u202F\u205F\u3000]
Example:

"https://your-cdn.com/samples/my-voice.wav"

preferred_name
string
required

Voice name prefix

Constraints:

  • 1–10 English letters or digits; underscores and other symbols are not supported
  • Uppercase letters are converted to lowercase
  • Does not need to be unique

Full generated name: qwen-audio-3.1-tts-flash-{preferred_name}-{32-character-id} for cloning, or qwen-audio-3.1-tts-flash-vd-{preferred_name}-{32-character-id} for design

For example, myvoice produces a cloned name such as qwen-audio-3.1-tts-flash-myvoice-5996beec833d41f4982158347ba97fae

Maximum string length: 10
Pattern: ^[a-zA-Z0-9]+$
Example:

"myvoice"

language
enum<string>

For cloning, the spoken language of the recording helps the model extract the voice more accurately. For design, this sets the voice's language preference; use the same language as preview_text

Defaults to zh if omitted

Available options:
zh,
en,
ja,
ko,
de,
fr,
it,
ru,
pt,
es
Example:

"zh"

target_model
enum<string>
default:qwen-audio-3.1-tts-flash

TTS model that will use the voice. Only one value is currently supported, and it is used if omitted; other values return 400

Available options:
qwen-audio-3.1-tts-flash
Example:

"qwen-audio-3.1-tts-flash"

callback_url
string<uri>

HTTPS callback URL for the task result

Timing:

  • Triggered when the task completes (completed) or fails (failed)
  • Sent after billing has been confirmed

Security requirements:

  • HTTPS only
  • Private-network IPs are prohibited (127.0.0.1, 10.x.x.x, 172.16–31.x.x, 192.168.x.x, etc.)
  • URL length at most 2048 characters

Delivery:

  • Timeout: 10 seconds
  • At most 3 retries after failure, with delays of 1 / 2 / 4 seconds
  • The callback body has the same format as the task query response
  • A 2xx status is success; other statuses trigger retries
Example:

"https://your-domain.com/webhooks/voice-completed"

Response

Voice creation task accepted

created
integer

Task creation timestamp

Example:

1775123456

id
string

Task ID

Example:

"task-unified-1775123456-abcd1234"

model
string

Model actually used

Example:

"voice-enrollment"

object
enum<string>

Specific task object type

Available options:
audio.generation.task
progress
integer

Task progress percentage (0–100)

Required range: 0 <= x <= 100
Example:

0

status
enum<string>

Task status

Available options:
pending,
processing,
completed,
failed
Example:

"pending"

task_info
object

Audio task details

type
enum<string>

Task output type

Available options:
audio
Example:

"audio"

usage
object

Usage and billing information