GPT Image 2.5 Flare & Sunburst are live on EvoLinkTry GPT Image 2.5

Qwen Audio 3.1 TTS Flash API

Use Qwen Audio 3.1 TTS Flash through EvoLink’s unified API: 68 voices, voice cloning and design. Check model IDs, token pricing and async integration.

AlibabaText-to-SpeechAvailable
From $1.765 / 1M output tokens
Text-to-SpeechSystem VoicesVoice CloningVoice DesignEmotion ControlMultilingual & Dialects
Production routeLive
Model Highlights
Natural speech with instruction-based emotion, dialect and pace control
Use Cases
Voice assistants, audiobooks, video narration, customer service
Input
Text up to 5,000 characters, plus an optional style instruction
Output
MP3, WAV or Opus audio with a system, cloned or designed voice

Try Qwen Audio 3.1 TTS Flash before you integrate the API.

Turn text into speech with a system voice, or switch the model to Voice Enrollment to clone or design your own voice and use it right away.

Create speech task

Fixed to qwen-audio-3.1-tts-flash
80/5,000

Up to 5,000 characters. Add tags such as [excited] or [laughing] in the text to shape the delivery.

Pick a system voice from the list, or type or paste a voice name here, including one you created. Voice names are case-sensitive. Pick a voice that supports the language of your text.

0/100

Optional. Controls emotion, pace or dialect. Up to 100 billing characters (Han characters, including Japanese kanji, count as 2; kana, Hangul and all other characters count as 1).

Optional hint for the language of the text.

1.0x
1.0x

Pitch also changes the duration, and has no effect when speed is not 1.0.

50
Estimate ~$0.00090.0578 cr
106 input + 468 output tokens are reserved up front, estimated from your text. You are billed for the tokens actually used: any excess is returned, and a shortfall is charged.
Ready
Generated audio

Your audio will appear here.

History
Qwen Audio 3.1 TTS Flash tasks will appear here.

Know your Qwen Audio 3.1 TTS Flash cost before you add credits.

Speech synthesis is billed by the tokens actually used, with separate input and output rates. Creating a custom voice is a small flat fee per voice. Failed tasks are refunded in full.

First Test Cost

One short sentence · 14 input + 33 output tokens
Measured sample
Credits0.0043

One short sentence

Approx. Cost~$0.0001

About 158,139 sentences like this with $10 in credits.

In Playground, keep the default voice and enter one sentence. Credits are reserved from the text length when the task is created, then settled to the tokens actually used: any excess is returned, and a shortfall is charged.

Qwen Audio 3.1 TTS Flash testing budget guide

Pick an amount based on how many tests you expect.
Add Credits
$10
About 475 long articles of 5,000 Chinese characters

Good for a first validation.

$50
About 2,375 long articles of 5,000 Chinese characters

Good for narrating a batch of articles or lessons.

$100
About 4,751 long articles of 5,000 Chinese characters

Good before production integration.

Qwen Audio 3.1 TTS Flash cost examples

One short sentence14 input + 33 output tokens · measured
~$0.0001
~0.0043 cr
A long article5,000 Chinese characters · 5,000 input + 11,300 output tokens · measured
~$0.022
~1.431 cr
Create a custom voice1 voice · voice cloning or voice design
$0.0001
0.001 cr

Token counts are measured samples. Actual usage depends on the language, the voice and how the text is read: digits, upper-case letters and symbols are read out one by one and use more output tokens.

Qwen Audio 3.1 TTS Flash pricing

Qwen Audio 3.1 TTS Flashqwen-audio-3.1-tts-flash
Output tokens$1.765/1M tokens

The generated speech. Longer audio uses more output tokens.

Input tokens$0.221/1M tokens

The text you send. The style instruction is not counted.

Voice Enrollmentvoice-enrollment
Create a voice$0.0001/voice

Voice cloning (audio_url) or voice design (voice_prompt), same price. The voice is bound to qwen-audio-3.1-tts-flash.

Billing notes

  • Speech synthesis reserves credits when the task is created, estimated from the text length. When the task completes, you are billed for the input and output tokens actually used: any excess is returned, and a shortfall is charged.
  • Each of the two token charges is rounded up to the smallest credit unit (0.0001 credits) before they are added.
  • Creating a voice is billed once. Synthesizing speech with a custom voice costs the same as with a system voice.
  • Failed tasks are refunded in full. Requests that fail validation return 400 and are not charged.

Qwen Audio 3.1 TTS Flash API: text-to-speech with system, cloned and designed voices

Qwen Audio 3.1 TTS Flash turns text into natural speech with 68 system voices and instruction-based control over emotion, pace and dialect. On the same page, Voice Enrollment creates your own voice for it: clone a consenting speaker from a short recording, or design a new voice from a text description. EvoLink exposes both as asynchronous tasks on one audio endpoint with one API key.

Input
Text ≤5,000 characters
Output
MP3 / WAV / Opus
Voices
68 system + your own
Sample rate
8–48 kHz
Billing
Per token used

Qwen Audio 3.1 TTS Flash model IDs for API calls

Qwen Audio 3.1 TTS Flash

qwen-audio-3.1-tts-flash

Text-to-speech. Send prompt and an optional voice; get an audio file back. Billed by input and output tokens.

Voice Enrollment

voice-enrollment

Create a custom voice for Qwen Audio 3.1 TTS Flash. Send audio_url to clone a voice from a recording, or voice_prompt and preview_text to design one from a description. Flat fee per voice.

Key Qwen Audio 3.1 TTS Flash API parameters

Speech synthesis parameters first, then the fields used by voice-enrollment. Open the API tab for the full request and response schema.

promptstring

Default: required

Text to synthesize, up to 5,000 characters. Emotion tags such as [excited] can be written inline.

voicestring

Default: longanhuan_v3.1

A system voice (case-sensitive), or a voice your account created with voice-enrollment.

instructionstring

Default: none

Natural-language control of emotion, pace or dialect. Up to 100 billing characters.

response_formatenum

Default: mp3

mp3, wav or opus. Opus supports 8, 12, 16, 24 and 48 kHz only.

speech_rate / pitchnumber

Default: 1.0

Both range from 0.5 to 2.0. Pitch also changes the duration and is ignored when speech_rate is not 1.0.

audio_urlstring · voice-enrollment

Default: required for cloning

Public HTTP(S) link to a WAV, MP3 or M4A sample with clear speech, up to 60 s and 10 MB. Selects voice cloning.

voice_prompt + preview_textstring · voice-enrollment

Default: required for design

Voice description up to 500 characters, and a 15–200 character sentence for the preview clip. Selects voice design.

preferred_namestring · voice-enrollment

Default: required

1–10 English letters or digits. It becomes part of the returned voice name.

Two ways to use Qwen Audio 3.1 TTS Flash: EvoLink API or Agent

Use the EvoLink API for product integration and batch jobs, or call it from Codex, Claude, or Gemini for fast creative and development workflows. Both paths share the same EvoLink API key, balance, model routes, and task history.

Option 1

Integrate with the EvoLink API

Best for: product backends, batch jobs, automated workflows

Call EvoLink’s unified voice API from your server and control the model ID, parameters, task queue, callbacks, and result storage.

  1. 1Validate output and cost with a real brief in Playground
  2. 2Create an EvoLink API key in the console
  3. 3Pick the model ID that matches your task from the route cards above
  4. 4Submit the task and retrieve the result by polling or HTTPS callback
Option 2

Call it with an Agent

Best for: creative and development tasks in Codex, Claude, and Gemini

Give the Agent your output goal, assets, and acceptance criteria. It can choose the route, assemble the request, track the task, and return the result without requiring you to hand-code every step.

  1. 1Set EVOLINK_API_KEY in your local environment; never put it in code or a prompt
  2. 2Describe the output goal, references, and the key parameters
  3. 3Ask the Agent to call a Qwen Audio 3.1 TTS Flash route and save the task ID
  4. 4Let the Agent poll the final state, download the result, and report failure reasons

What you can build with Qwen Audio 3.1 TTS Flash

Voice assistants with Qwen Audio 3.1 TTS Flash

Give an assistant a consistent voice and steer its tone per reply with a short style instruction.

Audiobooks and courses with Qwen Audio 3.1 TTS Flash

Narrate chapters up to 5,000 characters per request with a calm narrator voice.

Video narration and dubbing

Generate voice-over in Mandarin, Chinese dialects, English and more, with inline emotion tags.

A brand voice with Voice Enrollment

Clone a consenting speaker or design a new voice, then use the voice name in your synthesis calls for the next 6 hours.

Qwen Audio 3.1 TTS Flash API code example and error handling

This example shows the shortest runnable flow: create a task, save the returned task ID, then poll or wait for an HTTPS callback. Open the API tab for the complete parameter and response reference.

View complete API docs
cURL
curl -X POST https://api.evolink.ai/v1/audios/generations \
  -H "Authorization: Bearer $EVOLINK_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "qwen-audio-3.1-tts-flash",
    "prompt": "Hello, and welcome to EvoLink. This is a quick demo of Qwen Audio 3.1 TTS Flash.",
    "voice": "Emily_v3.1",
    "instruction": "Speak warmly, at a relaxed pace",
    "response_format": "mp3"
  }'

# Save the returned task ID, then query; the audio link is in results[0]:
curl https://api.evolink.ai/v1/tasks/{task_id} \
  -H "Authorization: Bearer $EVOLINK_API_KEY"

Parameter validation failed

Check every value against the parameter cards above: ranges, allowed options, asset counts, and file types must match the selected route.

Authentication or balance issue

Check the Authorization bearer token and confirm the available balance and task billing in the console.

Content or asset rejected

Review likeness rights, brand assets, readable text, sensitive content, and input-asset compliance.

Task failed or timed out

A Qwen Audio 3.1 TTS Flash task that reaches a final failed state is not charged. Keep the task ID, inspect the final state, and identify whether the cause is validation, review, or execution before retrying.

Callback not received

callback_url must be a public HTTPS URL, not localhost or a private network address. Always keep the task ID as a polling fallback.

System voices vs voice cloning vs voice design

FeatureSystem voiceVoice cloningVoice design
Model IDqwen-audio-3.1-tts-flashvoice-enrollmentvoice-enrollment
What you provideA voice name from the listA speech sample (audio_url)A voice description + preview text
What you getAudioA voice nameA voice name + preview clip
Setup costNone~$0.0001 per voice~$0.0001 per voice
Best forGetting started quicklyReproducing a specific, consenting speakerNew voices, characters, brand tone exploration

Qwen Audio 3.1 TTS Flash routing, task lifecycle, and billing

Two model IDs on one endpoint

Send qwen-audio-3.1-tts-flash to synthesize speech and voice-enrollment to create a voice. Both go to POST /v1/audios/generations and return a task ID.

Asynchronous tasks

Poll /v1/tasks/{task_id} or use callback_url for the result. Completion time varies with the text and voice operation; no fixed latency is guaranteed. Audio links expire after 24 hours.

Your voices stay yours

A custom voice only works with qwen-audio-3.1-tts-flash and only for the account that created it. Using someone else's voice name returns 404. A voice can be used for 6 hours after it is created.

Why use Qwen Audio 3.1 TTS Flash through EvoLink

Early validation

Unknown parameters, an invalid voice name or over-length text fail fast with 400 instead of reserving credits and failing later.

Unified API

One API key works across Qwen, Seed Audio, video, image, and LLM models.

Unified Balance

Track credits, spending, and model usage in one console.

Task Status Tracking

See whether a task is submitted, processing, completed, or failed, with the real failure reason.

Pay for what you use

Speech is settled on actual token usage, and credits reserved for a failed task are returned in full.

Other audio models on EvoLink besides Qwen Audio 3.1 TTS Flash

Doubao Seed Audio 1.0

Doubao Seed Audio 1.0

BytePlus prompt-based audio for voice, dialogue, sound effects, music, and ambience.

View API
Suno

Suno

Suno music generation with vocals, lyrics, and instrumental tracks.

View API

Qwen Audio 3.1 TTS Flash FAQ

Does EvoLink support streaming or WebSocket for this model?

This EvoLink route uses asynchronous HTTP tasks, without SSE or WebSocket streaming. Submit to POST /v1/audios/generations, then poll GET /v1/tasks/{task_id} for the audio URL. The provider’s streaming APIs use a different protocol.

Is Qwen Audio 3.1 TTS Flash the same as Qwen3-TTS?

No. This page covers the hosted qwen-audio-3.1-tts-flash model. Qwen3-TTS, Qwen Audio TTS-Next and Qwen Audio Realtime are separate models with different IDs, parameters, weights and streaming capabilities.

What is the Qwen Audio 3.1 TTS Flash model ID?

Use qwen-audio-3.1-tts-flash for speech synthesis and voice-enrollment to create a custom voice. Both are sent to POST /v1/audios/generations.

How much does Qwen Audio 3.1 TTS Flash cost?

Speech is billed by the tokens actually used, with separate input and output rates shown live in the Pricing section. Creating a custom voice is a small flat fee per voice. Failed tasks are refunded in full.

Why is more reserved than I am finally charged?

The token count is only known after synthesis, so credits are reserved from the text length when the task is created, usually more than the final charge. When it completes, you are billed for the tokens actually used and the excess is returned. Text that is read out character by character, such as strings of digits, letters or symbols, can use more than was reserved; the shortfall is then charged.

How do I create and use my own voice?

Call voice-enrollment with audio_url to clone a voice or with voice_prompt and preview_text to design one, read result_data.voice from the finished task, and pass it as the voice parameter to qwen-audio-3.1-tts-flash. A custom voice can be used for 6 hours after it is created; after that, create a new one.

What is the difference between voice cloning and voice design?

Voice cloning reproduces a real speaker from a short recording and returns the voice name only. Voice design creates a new voice from a text description and also returns a preview clip. Both cost the same and both work only with qwen-audio-3.1-tts-flash.

What makes a good voice cloning sample?

Required: WAV, MP3 or M4A, up to 60 s and 10 MB, 16 kHz or higher, with clear speech. For best results, use 10–20 seconds from one speaker, mono audio, pauses no longer than 2 seconds, and no background music or noise.

Can someone else use my custom voice?

No. A custom voice can only be used by the account that created it; other accounts get a 404 for the same voice name.

Whose voice can I clone?

Only your own voice or a voice you have explicit permission to use. Cloning someone's voice without consent is not allowed.

How do I control emotion, pace or dialect?

Pass a short instruction such as "speak slowly and gently", or write tags like [excited] directly in the text. The four multilingual voices also speak several Chinese dialects when the instruction asks for one.

What if a task fails?

Failed tasks are not charged—the reserved credits are returned in full. The task result shows the failure reason.