curl --request POST \
--url https://api.evolink.ai/v1/audios/generations \
--header 'Authorization: Bearer <token>' \
--header 'Content-Type: application/json' \
--data '
{
"model": "voice-enrollment",
"audio_url": "https://your-cdn.com/samples/my-voice.wav",
"preferred_name": "myvoice"
}
'{
"created": 1775123456,
"id": "task-unified-1775123456-abcd1234",
"model": "voice-enrollment",
"object": "audio.generation.task",
"progress": 0,
"status": "pending",
"task_info": {
"can_cancel": false,
"estimated_time": 60,
"audio_type": "voice_clone"
},
"type": "audio",
"usage": {
"credits_reserved": 0.001
}
}Voice Enrollment: Create a Custom Voice
- Create custom voices for Qwen Audio 3.1 TTS Flash with model
voice-enrollment. Two modes share the same model:audio_urlselects voice cloning, whilevoice_prompt+preview_textselects voice design. Supplying both mode selectors or neither returns400 - Voices created in either mode work only with
qwen-audio-3.1-tts-flashand only for the account that created them. Voices from the oldqwen-voice-designmodel cannot be used with this TTS; create new voices through this endpoint when migrating - Voice design also returns preview audio (fixed 24 kHz WAV); voice cloning returns only the voice name
- Asynchronous processing; use the returned task ID to query the result. Cloning takes about 15–30 seconds. Design waits for the full preview text to be synthesized: about 12 seconds for 30 characters and 49 seconds for 200 characters
- Charged per request, with the same price for cloning and design; failed tasks receive a full refund. Preview audio links are valid for 24 hours; save them promptly
- Only clone your own voice or a voice you have explicit permission to clone
Workflow:
- Call this endpoint with
audio_url(cloning) orvoice_prompt+preview_text(design), pluspreferred_name - Poll the task result to obtain
result_data.voice(the voice name) - Call Qwen Audio 3.1 TTS Flash and pass the full name as
voice
Task result (when status is completed):
| Field | Voice cloning | Voice design |
|---|---|---|
result_data.voice | qwen-audio-3.1-tts-flash-{preferred_name}-{32-character-id} | qwen-audio-3.1-tts-flash-vd-{preferred_name}-{32-character-id} (includes an extra vd- segment) |
result_data.voice_type | voice_clone | voice_design |
result_data.target_model | qwen-audio-3.1-tts-flash | qwen-audio-3.1-tts-flash |
results / result_data.preview_audio_url | Not returned | Preview audio URL (valid for 24 hours), plus sample_rate: 24000 and response_format: "wav" |
If design preview audio is occasionally unavailable, the task still completes (the voice has been created and billed), with preview_audio_unavailable: true and preview_audio_warning instead.
Voice lifetime:
- Custom voices created by cloning or design expire by default
6 hoursafter the creation task completes. Synthesis calls do not extend this lifetime - Synthesis after expiry returns
404(voice_expired). Create a new voice with this endpoint and use the new voice for synthesis
Voice slots: The upstream account has a cap on custom voices across the Qwen-Audio-TTS series (officially 1000, shared by cloning and design). Creation fails with a full refund when no slots remain.
Text length is measured in characters: Chinese characters, English characters, and punctuation each count as 1.
curl --request POST \
--url https://api.evolink.ai/v1/audios/generations \
--header 'Authorization: Bearer <token>' \
--header 'Content-Type: application/json' \
--data '
{
"model": "voice-enrollment",
"audio_url": "https://your-cdn.com/samples/my-voice.wav",
"preferred_name": "myvoice"
}
'{
"created": 1775123456,
"id": "task-unified-1775123456-abcd1234",
"model": "voice-enrollment",
"object": "audio.generation.task",
"progress": 0,
"status": "pending",
"task_info": {
"can_cancel": false,
"estimated_time": 60,
"audio_type": "voice_clone"
},
"type": "audio",
"usage": {
"credits_reserved": 0.001
}
}Authorizations
Bearer token authentication is required for all endpoints
Get an API key:
Visit API Key Management to obtain your API key
Add this request header:
Authorization: Bearer YOUR_API_KEY
Body
- Voice cloning (audio_url)
- Voice design (voice_prompt)
Create a voice from a recording of a person. Providing audio_url selects voice cloning; omit the design fields voice_prompt, preview_text, sample_rate, and response_format. Non-empty design text, a non-zero sample rate, or a non-empty format returns 400; sample_rate: 0/null and response_format: ""/null are treated as omitted.
Model name
voice-enrollment "voice-enrollment"
URL of the recording to clone. This field selects voice cloning and cannot be supplied together with voice_prompt
URL requirements:
- HTTP or HTTPS, publicly accessible without authentication
- Local or private-network addresses return
400(invalid_media_url) - Maximum
2048characters - Non-HTTP(S) protocols such as FTP return
400(invalid_media_url) - URLs only, not Base64; a Base64 data URI returns
400(invalid_parameter)
Audio requirements (the task may fail if unmet):
- WAV, MP3, or M4A
- At most 60 seconds; insufficient voiced content can also fail (a 2-second recording failed in testing)
- File size at most
10 MB - Sample rate of at least
16 kHz - Must contain clear human speech; silence, music, and other non-speech audio are rejected
Recording recommendations:
- 10–20 seconds, including at least 5 seconds of continuous, clear reading, with pauses no longer than 2 seconds
- Mono; for stereo recordings, only the first channel is used
- No background music, noise, or other voices; speak normally, do not sing
If audio cannot be downloaded or does not meet the requirements, the task fails and credits are fully refunded
Only clone your own voice or a voice you have explicit permission to clone
2048[^\u0009-\u000D\u0020\u0085\u00A0\u1680\u2000-\u200A\u2028\u2029\u202F\u205F\u3000]"https://your-cdn.com/samples/my-voice.wav"
Voice name prefix
Constraints:
- 1–10 English letters or digits; underscores and other symbols are not supported
- Uppercase letters are converted to lowercase
- Does not need to be unique
Full generated name: qwen-audio-3.1-tts-flash-{preferred_name}-{32-character-id} for cloning, or qwen-audio-3.1-tts-flash-vd-{preferred_name}-{32-character-id} for design
For example, myvoice produces a cloned name such as qwen-audio-3.1-tts-flash-myvoice-5996beec833d41f4982158347ba97fae
10^[a-zA-Z0-9]+$"myvoice"
For cloning, the spoken language of the recording helps the model extract the voice more accurately. For design, this sets the voice's language preference; use the same language as preview_text
Defaults to zh if omitted
zh, en, ja, ko, de, fr, it, ru, pt, es "zh"
TTS model that will use the voice. Only one value is currently supported, and it is used if omitted; other values return 400
| Value | Description |
|---|---|
qwen-audio-3.1-tts-flash | Qwen Audio 3.1 TTS Flash, non-streaming (default and only value) |
qwen-audio-3.1-tts-flash "qwen-audio-3.1-tts-flash"
HTTPS callback URL for the task result
Timing:
- Triggered when the task completes (
completed) or fails (failed) - Sent after billing has been confirmed
Security requirements:
- HTTPS only
- Private-network IPs are prohibited (127.0.0.1, 10.x.x.x, 172.16–31.x.x, 192.168.x.x, etc.)
- URL length at most
2048characters
Delivery:
- Timeout:
10seconds - At most
3retries after failure, with delays of1/2/4seconds - The callback body has the same format as the task query response
- A 2xx status is success; other statuses trigger retries
"https://your-domain.com/webhooks/voice-completed"
Response
Voice creation task accepted
Task creation timestamp
1775123456
Task ID
"task-unified-1775123456-abcd1234"
Model actually used
"voice-enrollment"
Specific task object type
audio.generation.task Task progress percentage (0–100)
0 <= x <= 1000
Task status
pending, processing, completed, failed "pending"
Audio task details
Show child attributes
Show child attributes
Task output type
audio "audio"
Usage and billing information
Show child attributes
Show child attributes