> ## Documentation Index
> Fetch the complete documentation index at: https://evolink.ai/docs/llms.txt
> Use this file to discover all available pages before exploring further.

# Qwen Audio 3.1 TTS Flash Speech Synthesis

> - Convert text to speech, up to `5000` characters per request
- Choose from 68 system voices; see the [voice list](/en/api-manual/audio-series/qwen-audio-tts/qwen-audio-3.1-tts-flash-voices). You can also use a voice you created with [Voice Enrollment](/en/api-manual/audio-series/qwen-audio-tts/voice-enrollment), through cloning or design
- Custom voices expire by default `6 hours` after the creation task completes; synthesis after expiry returns `404` (`voice_expired`). Create a new voice with Voice Enrollment
- Supports natural-language instructions (`instruction`), inline emotion and paralinguistic tags, SSML, and custom pronunciation (`hot_fix`)
- Only the parameters listed below are accepted; other parameters return `400` (`unsupported_parameter`)
- Asynchronous processing; use the returned task ID to [query the result](/en/api-manual/task-management/get-task-detail)
- Tasks cannot be canceled after submission
- Generated audio links are valid for 24 hours; save them promptly

**Billing:**
- Billed by actual token usage: input tokens (related to text length) and output tokens (related to generated audio duration) are priced separately
- Credits are reserved at submission based on text length (`usage.credits_reserved`). On completion, billing is settled against actual usage: excess is refunded and any shortfall is charged. Failed tasks receive a full refund
- Numbers, letters, and symbols may be read individually, producing much longer audio and more output tokens than ordinary prose of the same length; actual usage may exceed the reservation
- For the same text, instructions or tags that slow delivery (such as `[very slowly]`) can substantially increase output tokens. Adjusting `speech_rate` and inserting SSML pauses do not increase output tokens

**Task result (when `status` is `completed`):**

| Field | Description |
|---|---|
| `results[0]` | Audio URL |
| `result_data[0].audio_url` | Audio URL, identical to `results[0]` |
| `result_data[0].format` | Audio format |
| `result_data[0].sample_rate` | Sample rate (Hz) |
| `usage.input_tokens` / `usage.output_tokens` / `usage.total_tokens` | Token usage for this synthesis |
| `usage.credits_used` | Actual credits used |



## OpenAPI

````yaml en/api-manual/audio-series/qwen-audio-tts/qwen-audio-3.1-tts-flash.json POST /v1/audios/generations
openapi: 3.1.0
info:
  title: Qwen Audio 3.1 TTS Flash Speech Synthesis API
  description: >-
    Convert text to speech with 68 system voices, natural-language instructions,
    emotion tags, SSML, and custom pronunciation. Billing is based on actual
    token usage.
  license:
    name: MIT
  version: 1.0.0
servers:
  - url: https://api.evolink.ai
    description: Production environment
security:
  - bearerAuth: []
tags:
  - name: Speech Synthesis
    description: Qwen Audio 3.1 TTS Flash speech synthesis endpoints
paths:
  /v1/audios/generations:
    post:
      tags:
        - Speech Synthesis
      summary: Qwen Audio 3.1 TTS Flash Speech Synthesis
      description: >-
        - Convert text to speech, up to `5000` characters per request

        - Choose from 68 system voices; see the [voice
        list](/en/api-manual/audio-series/qwen-audio-tts/qwen-audio-3.1-tts-flash-voices).
        You can also use a voice you created with [Voice
        Enrollment](/en/api-manual/audio-series/qwen-audio-tts/voice-enrollment),
        through cloning or design

        - Custom voices expire by default `6 hours` after the creation task
        completes; synthesis after expiry returns `404` (`voice_expired`).
        Create a new voice with Voice Enrollment

        - Supports natural-language instructions (`instruction`), inline emotion
        and paralinguistic tags, SSML, and custom pronunciation (`hot_fix`)

        - Only the parameters listed below are accepted; other parameters return
        `400` (`unsupported_parameter`)

        - Asynchronous processing; use the returned task ID to [query the
        result](/en/api-manual/task-management/get-task-detail)

        - Tasks cannot be canceled after submission

        - Generated audio links are valid for 24 hours; save them promptly


        **Billing:**

        - Billed by actual token usage: input tokens (related to text length)
        and output tokens (related to generated audio duration) are priced
        separately

        - Credits are reserved at submission based on text length
        (`usage.credits_reserved`). On completion, billing is settled against
        actual usage: excess is refunded and any shortfall is charged. Failed
        tasks receive a full refund

        - Numbers, letters, and symbols may be read individually, producing much
        longer audio and more output tokens than ordinary prose of the same
        length; actual usage may exceed the reservation

        - For the same text, instructions or tags that slow delivery (such as
        `[very slowly]`) can substantially increase output tokens. Adjusting
        `speech_rate` and inserting SSML pauses do not increase output tokens


        **Task result (when `status` is `completed`):**


        | Field | Description |

        |---|---|

        | `results[0]` | Audio URL |

        | `result_data[0].audio_url` | Audio URL, identical to `results[0]` |

        | `result_data[0].format` | Audio format |

        | `result_data[0].sample_rate` | Sample rate (Hz) |

        | `usage.input_tokens` / `usage.output_tokens` / `usage.total_tokens` |
        Token usage for this synthesis |

        | `usage.credits_used` | Actual credits used |
      operationId: createQwenAudio31TtsFlash
      requestBody:
        required: true
        content:
          application/json:
            schema:
              $ref: '#/components/schemas/QwenAudioTtsRequest'
            examples:
              basic:
                summary: Minimal request (default voice)
                value:
                  model: qwen-audio-3.1-tts-flash
                  prompt: 我家的后面有一个很大的花园。
              aliases:
                summary: Text and format aliases
                value:
                  model: qwen-audio-3.1-tts-flash
                  input: 我家的后面有一个很大的花园。
                  format: mp3
              with_voice:
                summary: Select a voice and output format
                value:
                  model: qwen-audio-3.1-tts-flash
                  prompt: 各位听众朋友，大家好，欢迎收听晚间新闻。
                  voice: xuyanchu_v3.1
                  response_format: wav
                  sample_rate: 48000
              expressive:
                summary: Instructions and emotion tags
                value:
                  model: qwen-audio-3.1-tts-flash
                  prompt: '[excited]今天的天气真不错！[laughing]我们一起出去玩吧！'
                  voice: longanhuan_v3.1
                  instruction: 用欢快、热情的语气说
              english:
                summary: English voice
                value:
                  model: qwen-audio-3.1-tts-flash
                  prompt: Hello, this is 110. Please leave a message after the tone.
                  voice: Emily_v3.1
                  language: en
                  speech_rate: 0.9
              ssml:
                summary: SSML pauses
                value:
                  model: qwen-audio-3.1-tts-flash
                  prompt: <speak>欢迎收听今天的节目。<break time="1s"/>我们马上开始。</speak>
                  enable_ssml: true
              full_params:
                summary: Full parameters
                value:
                  model: qwen-audio-3.1-tts-flash
                  prompt: 今天的天气真不错，适合出去走走。
                  voice: yuxiaoyun_v3.1
                  response_format: mp3
                  sample_rate: 24000
                  volume: 60
                  speech_rate: 1.1
                  pitch: 1
                  instruction: 语气轻松自然，像在和朋友聊天
                  language: zh
                  enable_ssml: false
                  hot_fix:
                    pronunciation:
                      - 天气: tian1 qi4
                    replace:
                      - 走走: 走一走
                  enable_aigc_tag: false
                  callback_url: https://your-domain.com/webhooks/tts-completed
      responses:
        '200':
          description: Speech synthesis task created successfully
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/QwenAudioTtsResponse'
        '400':
          description: Invalid request parameters
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/ErrorResponse'
              examples:
                missing_prompt:
                  summary: Missing prompt
                  value:
                    error:
                      code: missing_prompt
                      message: prompt is required for qwen-audio-3.1-tts-flash
                      type: invalid_request_error
                prompt_too_long:
                  summary: prompt exceeds 5000 characters
                  value:
                    error:
                      code: prompt_too_long
                      message: prompt must be at most 5000 characters, got 5210
                      type: invalid_request_error
                prompt_too_long_after_replace:
                  summary: Text exceeds 5000 characters after hot_fix.replace
                  value:
                    error:
                      code: prompt_too_long
                      message: >-
                        prompt must be at most 5000 characters after
                        hot_fix.replace is applied, got 5120
                      type: invalid_request_error
                hot_fix_too_many_entries:
                  summary: hot_fix exceeds 200 entries
                  value:
                    error:
                      code: invalid_parameter
                      message: >-
                        hot_fix may contain at most 200 entries in total, got
                        201
                      type: invalid_request_error
                invalid_voice:
                  summary: Voice is not in the supported list
                  value:
                    error:
                      code: invalid_voice
                      message: >-
                        voice "longanhuan" is not available for
                        qwen-audio-3.1-tts-flash; use a system voice from the
                        voice list or a voice you created with voice-enrollment
                      type: invalid_request_error
                instruction_too_long:
                  summary: instruction exceeds 100 billing characters
                  value:
                    error:
                      code: instruction_too_long
                      message: >-
                        instruction must be at most 100 billing characters (CJK
                        characters count as 2), got 124
                      type: invalid_request_error
                wrong_parameter_name:
                  summary: Incorrect parameter name (instructions)
                  value:
                    error:
                      code: unsupported_parameter
                      message: instructions is not supported; use "instruction"
                      type: invalid_request_error
                unsupported_parameter:
                  summary: Unsupported parameter supplied
                  value:
                    error:
                      code: unsupported_parameter
                      message: seed is not supported for qwen-audio-3.1-tts-flash
                      type: invalid_request_error
                opus_sample_rate:
                  summary: Unsupported sample rate for opus
                  value:
                    error:
                      code: invalid_parameter
                      message: >-
                        sample_rate 44100 is not supported with
                        response_format=opus; use one of: 8000, 12000, 16000,
                        24000, 48000
                      type: invalid_request_error
        '401':
          description: Unauthenticated, invalid or expired token
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/ErrorResponse'
              example:
                error:
                  code: unauthorized
                  message: Invalid or expired token
                  type: authentication_error
        '402':
          description: Insufficient quota; top up required
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/ErrorResponse'
              example:
                error:
                  code: insufficient_quota
                  message: Insufficient quota. Please top up your account.
                  type: insufficient_quota
        '403':
          description: Access denied
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/ErrorResponse'
              example:
                error:
                  code: model_access_denied
                  message: >-
                    Token does not have access to model:
                    qwen-audio-3.1-tts-flash
                  type: invalid_request_error
        '404':
          description: Voice does not exist or has expired
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/ErrorResponse'
              examples:
                voice_not_found:
                  summary: Voice does not exist or belongs to another account
                  value:
                    error:
                      code: voice_not_found
                      message: >-
                        voice
                        "qwen-audio-3.1-tts-flash-myvoice-5996beec833d41f4982158347ba97fae"
                        not found
                      type: invalid_request_error
                voice_expired:
                  summary: Custom voice has expired; create a new one
                  value:
                    error:
                      code: voice_expired
                      message: >-
                        voice
                        "qwen-audio-3.1-tts-flash-myvoice-5996beec833d41f4982158347ba97fae"
                        has expired; create a new one with voice-enrollment
                      type: invalid_request_error
        '429':
          description: Rate limit exceeded
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/ErrorResponse'
              example:
                error:
                  code: rate_limit_exceeded
                  message: Too many requests, please try again later
                  type: rate_limit_error
        '500':
          description: Internal server error
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/ErrorResponse'
              example:
                error:
                  code: internal_error
                  message: Internal server error
                  type: api_error
components:
  schemas:
    QwenAudioTtsRequest:
      type: object
      required:
        - model
      description: >-
        Provide at least one non-empty text field, `prompt` or `input`; if both
        are supplied, their contents must match. If both `response_format` and
        `format` are supplied, their values must match.
      anyOf:
        - required:
            - prompt
          properties:
            prompt:
              pattern: \S
        - required:
            - input
          properties:
            input:
              pattern: \S
      properties:
        model:
          type: string
          description: Model name
          enum:
            - qwen-audio-3.1-tts-flash
          default: qwen-audio-3.1-tts-flash
          example: qwen-audio-3.1-tts-flash
        prompt:
          type: string
          description: >-
            Text to synthesize


            **Constraints:**

            - Maximum `5000` characters

            - Keep punctuation in long text: long continuous passages without
            sentence breaks may be truncated upstream at about `1500` output
            tokens (about 120 seconds of audio). The task still reports success
            and is billed for the tokens actually generated; the gateway cannot
            detect this truncation

            - You can also use `input`. At least one non-empty text field is
            required; supplying just one is recommended. Different contents in
            both fields return `400` (`parameter_conflict`)

            - Text must be in a language supported by the selected voice;
            otherwise pronunciation may be incorrect


            **Emotion and paralinguistic tags:** Embed tags directly in the text
            without extra parameters; tag text counts toward billing characters

            - **Control tags**: Set the emotion or style of the following text
            until the next control tag. `[sad]` sad, `[amazed]` amazed, `[deep
            and loud shouting]` deep, loud shouting, `[trembling]` trembling,
            `[angry]` angry, `[excited]` excited, `[sarcastic]` sarcastic,
            `[curious]` curious, `[like dracula]` low and eerie, `[bored]`
            bored, `[tired]` tired, `[scornful]` scornful, `[shouting]`
            shouting, `[asmr]` soft ASMR whispers, `[panicked]` panicked,
            `[mischievously]` mischievous, `[empathetic]` empathetic,
            `[whispers]` whispering, `[reluctantly]` reluctant, `[crying]`
            crying, `[serious]` serious, `[very slowly]` very slowly, `[very
            fast]` very fast

            - **Paralinguistic tags**: Insert a vocal effect at this position
            without changing the emotion of the surrounding text. `[gasp]` gasp,
            `[sighing]` sigh, `[clears throat]` clear the throat, `[giggles]`
            giggle, `[laughing]` laugh, `[cough]` cough, `[snorts]` snort


            **Example:** `[excited]今天的天气真不错！[laughing]我们一起出去玩吧！`


            When `enable_ssml` is `true`, this field is parsed as SSML
          maxLength: 5000
          example: 我家的后面有一个很大的花园。
        input:
          type: string
          description: >-
            Alias for `prompt`, with the same length limits and usage rules


            - Provide at least one non-empty text field, `prompt` or `input`

            - If both are supplied, their contents must match; otherwise `400`
            (`parameter_conflict`) is returned
          maxLength: 5000
          example: 我家的后面有一个很大的花园。
        voice:
          type: string
          description: >-
            Voice name, case-sensitive


            - 68 system voices; see the [voice
            list](/en/api-manual/audio-series/qwen-audio-tts/qwen-audio-3.1-tts-flash-voices)
            for names, gender, and use cases

            - Defaults to `longanhuan_v3.1` if omitted

            - You can also supply a voice you created with [Voice
            Enrollment](/en/api-manual/audio-series/qwen-audio-tts/voice-enrollment):
            cloned voices follow
            `qwen-audio-3.1-tts-flash-{prefix}-{32-character-id}`, and designed
            voices follow
            `qwen-audio-3.1-tts-flash-vd-{prefix}-{32-character-id}`. Only the
            account that created the voice can use it. Voices from other models,
            such as `qwen-tts-vd-…` from `qwen-voice-design`, return `400`
            (`invalid_voice`). Missing voices or voices owned by another account
            return `404` (`voice_not_found`)

            - Custom voices expire by default `6 hours` after the creation task
            completes; synthesis after expiry returns `404` (`voice_expired`).
            Create a new voice with Voice Enrollment
          default: longanhuan_v3.1
          example: longanhuan_v3.1
        response_format:
          type: string
          description: >-
            Output audio format: `mp3`, `wav`, or `opus`; defaults to `mp3`


            - `opus` uses an Ogg Opus container

            - You can also use `format`. Supplying just one field is
            recommended; conflicting values return `400` (`parameter_conflict`)
          enum:
            - mp3
            - wav
            - opus
          default: mp3
          example: mp3
        format:
          type: string
          description: >-
            Alias for `response_format`; supports `mp3`, `wav`, and `opus`


            - Defaults to `mp3` when neither field is supplied

            - If both fields are supplied, their values must match; otherwise
            `400` (`parameter_conflict`) is returned
          enum:
            - mp3
            - wav
            - opus
          example: mp3
        sample_rate:
          type:
            - integer
            - 'null'
          description: >-
            Output sample rate (Hz)


            - `22050` and `44100` are not supported when `response_format` is
            `opus`

            - Omitted or `null` uses the default; `0` or a value outside the
            list returns `400`
          enum:
            - 8000
            - 12000
            - 16000
            - 22050
            - 24000
            - 44100
            - 48000
            - null
          default: 24000
          example: 24000
        volume:
          type: integer
          description: Volume, from `0` to `100`
          minimum: 0
          maximum: 100
          default: 50
          example: 50
        speech_rate:
          type: number
          description: >-
            Speech rate multiplier


            - `1.0`: normal speed (default)

            - `2.0`: double speed; `0.5`: half speed


            Range: `0.5` to `2.0`. Adjusting speech rate does not change output
            token count
          minimum: 0.5
          maximum: 2
          default: 1
          example: 1
        pitch:
          type: number
          description: >-
            Pitch multiplier


            - `1.0`: default pitch

            - Above `1.0` raises pitch; below `1.0` lowers it


            Range: `0.5` to `2.0`


            **Changing pitch also changes speech rate and audio duration**

            - Higher pitch speeds up speech and shortens audio; lower pitch
            slows it down and lengthens audio. Duration varies approximately
            inversely with the square of the pitch value

            - For a sentence lasting about 2.8 seconds at `1.0`: `0.8` is about
            4.3 seconds, `1.2` about 2.1 seconds, `0.5` about 10.9 seconds, and
            `2.0` about 0.7 seconds

            - Small adjustments between `0.8` and `1.2` are recommended; values
            near `0.5` or `2.0` make speech noticeably too slow or too fast

            - If `speech_rate` is also supplied with a value other than `1.0`,
            `pitch` has no effect; the two cannot be combined

            - Adjusting pitch does not change output token count
          minimum: 0.5
          maximum: 2
          default: 1
          example: 1
        instruction:
          type: string
          description: >-
            Natural-language instructions to control emotion, tone, character,
            dialect, and other delivery choices


            **Constraints:**

            - Maximum `100` billing characters: Han characters (including
            Japanese kanji and Korean hanja) count as 2; other characters count
            as 1, including kana and Hangul (about 50 Han characters or 100
            English characters). Exceeding the limit returns `400`


            **Examples:**

            - `用欢快、热情的语气说` (speak cheerfully and enthusiastically)

            - `请用上海话表达` (use Shanghainese; multilingual and dialect voices)

            - `Speak slowly in a calm and gentle tone`


            Instructions do not count toward input tokens, but can change the
            generated audio duration and therefore output tokens


            > The parameter is `instruction` (singular); `instructions` returns
            `400`
          example: 用欢快、热情的语气说
        language:
          type: string
          description: >-
            Target-language hint to improve the reading of numbers,
            abbreviations, and symbols, and synthesis in less common languages


            For example, with `zh`, `110` in `hello, this is 110` is read as
            “yao yao ling” in Chinese


            | Value | Language | Value | Language |

            |---|---|---|---|

            | `zh` | Chinese | `th` | Thai |

            | `en` | English | `id` | Indonesian |

            | `fr` | French | `vi` | Vietnamese |

            | `de` | German | `es` | Spanish |

            | `ja` | Japanese | `it` | Italian |

            | `ko` | Korean | `ms` | Malay |

            | `ru` | Russian | `fil` | Filipino |

            | `pt` | Portuguese | `ar` | Arabic |


            The model detects the language if omitted; this parameter does not
            translate text
          enum:
            - zh
            - en
            - fr
            - de
            - ja
            - ko
            - ru
            - pt
            - th
            - id
            - vi
            - es
            - it
            - ms
            - fil
            - ar
          example: zh
        enable_ssml:
          type: boolean
          description: >-
            Parse `prompt` as SSML


            When enabled, SSML tags can be used, such as `<break time="1s"/>`
            for a pause:

            `<speak>欢迎收听今天的节目。<break time="1s"/>我们马上开始。</speak>`


            Pauses inserted with SSML do not count toward output tokens
          default: false
          example: false
        hot_fix:
          type: object
          description: >-
            Custom pronunciation and text replacement to correct polyphonic
            characters, proper names, and other pronunciations


            - `pronunciation`: annotate words with pinyin (separate syllables
            with spaces and mark tones with digits, such as `tian1 qi4`) to
            override the default pronunciation

            - `replace`: replace specified words before synthesis. Synthesis and
            billing use the replaced text, which must also stay within `5000`
            characters; exceeding the limit returns `400` (`prompt_too_long`)


            The two lists may contain at most `200` entries in total, counted as
            key-value pairs within the objects. Exceeding the limit returns
            `400` (`invalid_parameter`)


            Provide at least one list. Each provided list must be a non-empty
            array of objects in `{"word": "value"}` form


            **Example:**

            ```json

            {
              "pronunciation": [{"天气": "tian1 qi4"}],
              "replace": [{"今天": "金天"}]
            }

            ```
          properties:
            pronunciation:
              type: array
              description: >-
                Custom pronunciation list; each item is an object such as
                `{"word": "pinyin"}`
              items:
                type: object
                additionalProperties:
                  type: string
              example:
                - 天气: tian1 qi4
            replace:
              type: array
              description: >-
                Text replacement list; each item is an object such as
                `{"original word": "replacement text"}`
              items:
                type: object
                additionalProperties:
                  type: string
              example:
                - 今天: 金天
        enable_aigc_tag:
          type: boolean
          description: >-
            Embed an invisible AIGC marker in generated audio (applies to `wav`
            / `mp3` / `opus`)
          default: false
          example: false
        callback_url:
          type: string
          description: >-
            HTTPS callback URL for the task result


            **Timing:**

            - Triggered when the task completes (`completed`) or fails
            (`failed`); this model does not support cancellation

            - Sent after billing has been confirmed


            **Security requirements:**

            - HTTPS only

            - Private-network IPs are prohibited (127.0.0.1, 10.x.x.x,
            172.16–31.x.x, 192.168.x.x, etc.)

            - URL length at most `2048` characters


            **Delivery:**

            - Timeout: `10` seconds

            - At most `3` retries after failure, with delays of `1` / `2` / `4`
            seconds

            - The callback body has the same format as the task query response

            - A 2xx status is success; other statuses trigger retries
          format: uri
          example: https://your-domain.com/webhooks/tts-completed
    QwenAudioTtsResponse:
      type: object
      properties:
        created:
          type: integer
          description: Task creation timestamp
          example: 1790000000
        id:
          type: string
          description: Task ID
          example: task-unified-1790000000-abcd1234
        model:
          type: string
          description: Model actually used
          example: qwen-audio-3.1-tts-flash
        object:
          type: string
          enum:
            - audio.generation.task
          description: Specific task object type
        progress:
          type: integer
          description: Task progress percentage (0–100)
          minimum: 0
          maximum: 100
          example: 0
        status:
          type: string
          description: Task status
          enum:
            - pending
            - processing
            - completed
            - failed
          example: pending
        task_info:
          $ref: '#/components/schemas/AudioTaskInfo'
          description: Audio task details
        type:
          type: string
          enum:
            - audio
          description: Task output type
          example: audio
        usage:
          $ref: '#/components/schemas/AudioUsage'
          description: Usage and billing information
    ErrorResponse:
      type: object
      properties:
        error:
          type: object
          properties:
            code:
              type: string
              description: Error code identifier
            message:
              type: string
              description: Error message
            type:
              type: string
              description: Error type
    AudioTaskInfo:
      type: object
      properties:
        can_cancel:
          type: boolean
          description: >-
            Whether the task can be canceled (this model does not support
            cancellation)
          example: false
        estimated_time:
          type: integer
          description: >-
            Estimated completion time in seconds; increases with text length, up
            to about `90` seconds
          minimum: 0
          example: 3
        audio_type:
          type: string
          description: Audio task type
          example: tts
    AudioUsage:
      type: object
      description: Usage information
      properties:
        credits_reserved:
          type: number
          description: >-
            Credits reserved based on text length; settled against actual token
            usage on completion, with excess refunded or shortfall charged
          minimum: 0
          example: 0.0144
  securitySchemes:
    bearerAuth:
      type: http
      scheme: bearer
      description: >-
        ## Bearer token authentication is required for all endpoints


        **Get an API key:**


        Visit [API Key Management](https://evolink.ai/dashboard/keys) to obtain
        your API key


        **Add this request header:**

        ```

        Authorization: Bearer YOUR_API_KEY

        ```

````


> ## Documentation Index
> Fetch the complete documentation index at: https://evolink.ai/docs/llms.txt
> Use this file to discover all available pages before exploring further.

# Query Task Status

> Query the status, progress, and result of a Qwen Audio 3.1 TTS Flash speech synthesis task by task ID



## OpenAPI

````yaml en/api-manual/task-management/get-task-detail.json GET /v1/tasks/{task_id}
openapi: 3.1.0
info:
  title: Get Task Details API
  description: >-
    Query the status, progress, and result information of asynchronous tasks by
    task ID
  license:
    name: MIT
  version: 1.0.0
servers:
  - url: https://api.evolink.ai
    description: Production environment
security:
  - bearerAuth: []
tags:
  - name: Task Management
    description: Asynchronous task management related APIs
paths:
  /v1/tasks/{task_id}:
    get:
      tags:
        - Task Management
      summary: Query Task Status
      description: >-
        Query the status, progress, and result information of asynchronous tasks
        by task ID
      operationId: getTaskDetail
      parameters:
        - name: task_id
          in: path
          required: true
          schema:
            type: string
          description: >-
            Task ID, ignore {} when querying, append the id from the async task
            response body at the end of the path
          example: task-unified-1756817821-4x3rx6ny
      responses:
        '200':
          description: Task status details
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/TaskDetailResponse'
              examples:
                completed:
                  summary: Completed
                  value:
                    created: 1790000000
                    duration: 3
                    id: task-unified-1790000000-abcd1234
                    model: qwen-audio-3.1-tts-flash
                    object: audio.generation.task
                    progress: 100
                    result_data:
                      - audio_url: https://example.com/speech.mp3
                        format: mp3
                        sample_rate: 24000
                    results:
                      - https://example.com/speech.mp3
                    status: completed
                    task_info:
                      can_cancel: false
                    type: audio
                    usage:
                      cost:
                        credits: 0.0043
                        usd: 0.0001
                        cny: 0.0005
                      credits_used: 0.0043
                      input_tokens: 14
                      output_tokens: 33
                      total_tokens: 47
                failed:
                  summary: Failed (reserved credits refunded in full)
                  value:
                    created: 1790000000
                    duration: 2
                    error:
                      code: invalid_prompt
                      message: The text could not be synthesized. Please check the input text.
                    id: task-unified-1790000000-abcd1234
                    model: qwen-audio-3.1-tts-flash
                    object: audio.generation.task
                    progress: 0
                    status: failed
                    task_info:
                      can_cancel: false
                    type: audio
        '401':
          description: 'Unauthenticated: API key missing, invalid, expired or disabled'
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/ErrorResponse'
              example:
                error:
                  code: unauthorized
                  message: 'Invalid API key (request id: 20260908103000123456789abcdef)'
                  param: null
                  type: authentication_error
        '403':
          description: Account disabled
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/ErrorResponse'
              example:
                error:
                  code: user_disabled
                  message: >-
                    Your account has been disabled (request id:
                    20260908103000123456789abcdef)
                  param: null
                  type: invalid_request_error
        '404':
          description: >-
            Task does not exist or does not belong to the current API key (also
            returned for malformed task IDs)
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/ErrorResponse'
              example:
                error:
                  message: Task not found
                  type: invalid_request_error
                  param: ''
                  code: task_not_found
        '429':
          description: Request rate limit exceeded (Retry-After header is also set)
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/ErrorResponse'
              example:
                error:
                  code: rate_limit_exceeded
                  message: >-
                    You have reached the request limit: max 60 requests per 1
                    minutes. Retry after 30 seconds. (request id:
                    20260908103000123456789abcdef)
                  param: null
                  retry_after: 30
                  retryable: true
                  type: rate_limit_error
        '500':
          description: Internal server error
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/ErrorResponse'
              example:
                error:
                  message: Failed to get task status
                  type: api_error
                  param: ''
                  code: internal_error
components:
  schemas:
    TaskDetailResponse:
      type: object
      properties:
        created:
          type: integer
          description: Task creation timestamp
          example: 1790000000
        id:
          type: string
          description: Task ID
          example: task-unified-1790000000-abcd1234
        model:
          type: string
          description: Model used
          example: qwen-audio-3.1-tts-flash
        object:
          type: string
          description: Task object type
          enum:
            - audio.generation.task
          example: audio.generation.task
        progress:
          type: integer
          minimum: 0
          maximum: 100
          description: Task progress percentage
          example: 100
        status:
          type: string
          description: >-
            Task status. Note: while the task is still pending upstream the
            query API already reports "processing", so clients normally only see
            processing / completed / failed.
          enum:
            - pending
            - processing
            - completed
            - failed
          example: completed
        type:
          type: string
          description: Task output type
          enum:
            - audio
          example: audio
        duration:
          type: integer
          description: >-
            Task duration in seconds. Only present when status is "completed" or
            "failed".
          example: 3
        results:
          type: array
          items:
            type: string
            format: uri
          description: >-
            Audio URL list with one item, valid for 24 hours. Only present when
            status is "completed".
          example:
            - https://example.com/speech.mp3
        result_data:
          type: array
          description: >-
            Result details with one item. Only present when status is
            "completed".
          items:
            type: object
            properties:
              audio_url:
                type: string
                format: uri
                description: Audio URL, identical to results[0]
                example: https://example.com/speech.mp3
              format:
                type: string
                description: Audio format
                enum:
                  - mp3
                  - wav
                  - opus
                example: mp3
              sample_rate:
                type: integer
                description: Sample rate (Hz)
                example: 24000
        usage:
          type: object
          description: >-
            Usage and billing information. Only present when status is
            "completed".
          properties:
            credits_used:
              type: number
              description: >-
                Credits actually charged for this task. Omitted only if billing
                could not be settled yet.
              example: 0.0043
            cost:
              type: object
              description: >-
                Cost of this task expressed in three currencies, converted from
                credits_used (values are rounded up).
              properties:
                credits:
                  type: number
                  description: Credits charged (same as credits_used)
                  example: 0.0043
                usd:
                  type: number
                  description: Equivalent amount in USD
                  example: 0.0001
                cny:
                  type: number
                  description: Equivalent amount in CNY
                  example: 0.0005
            input_tokens:
              type: integer
              description: Input tokens actually used (related to text length)
              example: 14
            output_tokens:
              type: integer
              description: >-
                Output tokens actually used (related to generated audio
                duration)
              example: 33
            total_tokens:
              type: integer
              description: input_tokens + output_tokens
              example: 47
        error:
          type: object
          nullable: true
          description: >-
            Error information when the task fails (only present when status is
            "failed"). Note: error.code here is a string-type business error
            code, different from HTTP status codes. Failed tasks are refunded in
            full.
          properties:
            code:
              type: string
              description: Business error code (string type)
              example: invalid_prompt
            message:
              type: string
              description: User-friendly error description with troubleshooting tips
              example: The text could not be synthesized. Please check the input text.
            details:
              type: string
              description: Additional error details, present only when available
        task_info:
          type: object
          description: Task detailed information
          properties:
            can_cancel:
              type: boolean
              description: >-
                Whether the task can be cancelled. Always false: this model does
                not support cancellation
              example: false
    ErrorResponse:
      type: object
      properties:
        error:
          type: object
          properties:
            code:
              type: string
              description: Error code (string type)
              example: task_not_found
            message:
              type: string
              description: >-
                Error description message. Errors raised by middleware (401,
                429) are suffixed with " (request id: ...)"
              example: Task not found
            type:
              type: string
              description: Error type
              example: invalid_request_error
            param:
              type: string
              nullable: true
              description: Related parameter name; null or empty for this endpoint
            retry_after:
              type: integer
              description: Seconds to wait before retrying (429 only)
              example: 30
            retryable:
              type: boolean
              description: Whether the request can be retried (429 only)
              example: true
  securitySchemes:
    bearerAuth:
      type: http
      scheme: bearer
      description: >-
        ##All APIs require Bearer Token authentication##


        **Get API Key:**


        Visit [API Key Management Page](https://evolink.ai/dashboard/keys) to
        get your API Key


        **Add to request header:**

        ```

        Authorization: Bearer YOUR_API_KEY

        ```

````


> ## Documentation Index
> Fetch the complete documentation index at: https://evolink.ai/docs/llms.txt
> Use this file to discover all available pages before exploring further.

# Qwen Audio 3.1 TTS Flash Voices

> IDs, gender, languages, and usage examples for 68 system voices, plus custom voice instructions.

Qwen Audio 3.1 TTS Flash selects a voice with the `voice` parameter in the [speech synthesis API](/docs/en/api-manual/audio-series/qwen-audio-tts/qwen-audio-3.1-tts-flash). There are **68** system voices, ready to use without creating a voice first.

* `voice` is **case-sensitive**; copy its value exactly as shown below
* If `voice` is omitted, the default is `longanhuan_v3.1`
* Use the table’s `voice` column in requests; names, gender, voice qualities, and use cases help you choose a voice
* The text must be in a language supported by the selected voice; otherwise pronunciation may be incorrect or speech unnatural

## Usage Examples

Send the following request body to `POST /v1/audios/generations`, replacing `voice` with the desired voice ID. See the [speech synthesis API](/docs/en/api-manual/audio-series/qwen-audio-tts/qwen-audio-3.1-tts-flash) for all parameters and task queries.

<CodeGroup>
  ```json Chinese voice theme={null}
  {
    "model": "qwen-audio-3.1-tts-flash",
    "prompt": "欢迎使用语音合成，让文字变成自然流畅的声音。",
    "voice": "longanhuan_v3.1"
  }
  ```

  ```json English voice theme={null}
  {
    "model": "qwen-audio-3.1-tts-flash",
    "prompt": "Welcome to our audio guide. Let us turn your words into speech.",
    "voice": "Emily_v3.1",
    "language": "en"
  }
  ```
</CodeGroup>

<Tip>
  All voices can use `instruction` and inline emotion tags, such as `[excited]` and `[laughing]`, to control emotion and tone. See the `prompt` and `instruction` descriptions in the [speech synthesis API](/docs/en/api-manual/audio-series/qwen-audio-tts/qwen-audio-3.1-tts-flash).
</Tip>

## Multilingual and Dialect Voices

**4** voices. Each supports all the dialects and languages below:

* **Dialects**: Shanghainese, Cantonese, Northeastern Chinese, Chongqing, Shaanxi, Yunnan, Ningbo, and Gansu dialects
* **Languages**: Japanese, Korean, French, German, Portuguese, Italian, Vietnamese, and Indonesian

For dialect synthesis, specify the dialect in `instruction`, for example `请用上海话表达` (use Shanghainese).

| Name | voice | Gender | Language |
| - | - | - | - |
| Long Anhuan (default) | `longanhuan_v3.1` | Female | Mandarin, dialects and multiple languages |
| Long Anlingxin | `longanlingxin_v3.1` | Female | Mandarin, dialects and multiple languages |
| Long Anfengyue | `longanfengyue_v3.1` | Female | Mandarin, dialects and multiple languages |
| Xu Nanchuan | `xunanchuan_v3.1` | Male | Mandarin, dialects and multiple languages |

## Premium Chinese Voices

**22** voices, supporting Mandarin Chinese only.

| Name | voice | Gender | Voice qualities | Use cases |
| - | - | - | - | - |
| Yu Xiaoyun | `yuxiaoyun_v3.1` | Female | energetic, friendly, natural | advertising and marketing, broadcasting, customer service assistant, narration |
| Qiao Xiaojiao | `qiaoxiaojiao_v3.1` | Female | charming, cute | advertising and marketing, customer service assistant, audiobooks |
| Xia Xiaochen | `xiaxiaochen_v3.1` | Female | energetic, bright | advertising and marketing, audiobooks |
| An Mingyuan | `anmingyuan_v3.1` | Male | clear, natural | advertising and marketing, audiobooks, narration |
| Wen Huaiqing | `wenhuaiqing_v3.1` | Female | clear, soft | children’s stories, customer service assistant, advertising and marketing, news reading |
| An Xiaolan | `anxiaolan_v3.1` | Female | sweet and clear, pure | audiobooks, customer service assistant, narration, news reading, advertising and marketing |
| Xie Shurou | `xieshurou_v3.1` | Female | soft, natural, articulate | audiobooks, customer service assistant, narration |
| Bai Qinglan | `baiqinglan_v3.1` | Female | bright, innocent | voice assistant, customer service assistant |
| Xu Yuyuan | `xuyuyuan_v3.1` | Female | articulate, mature, textured | advertising and marketing, news reading, narration, customer service assistant, audiobooks |
| An Ruorou | `anruorou_v3.1` | Female | breathy, articulate | narration, voice assistant |
| Wen Huaizhi | `wenhuaizhi_v3.1` | Female | steady, mature | audiobooks, news reading, advertising and marketing, customer service assistant, narration |
| Xiao Xingzhi | `xiaoxingzhi_v3.1` | Female | dignified, regal | news reading, audiobooks, narration, customer service assistant |
| Gu Yunshu | `guyunshu_v3.1` | Female | mature, steady | music radio, customer service assistant, audiobooks, narration |
| Huo Zhuoshi | `huozhuoshi_v3.1` | Male | clear | audiobooks, advertising and marketing, narration |
| Ye Qinghe | `yeqinghe_v3.1` | Female | friendly, gentle | audiobooks, advertising and marketing, narration, customer service assistant |
| Yun Huanhuan | `yunhuanhuan_v3.1` | Female | high-pitched, enthusiastic | audiobooks, narration, customer service assistant |
| Xu Xiaoqiao | `xuxiaoqiao_v3.1` | Female | natural, playful | audiobooks, narration, customer service assistant |
| Bai Anran | `baianran_v3.1` | Female | low-pitched, rich, breathy | dubbing and explanations, audiobooks, narration |
| Xu Yanchu | `xuyanchu_v3.1` | Female | composed, resonant | news reading, audiobooks |
| Ye Zhiqing | `yezhiqing_v3.1` | Female | lively, natural | children’s stories, customer service assistant, voice assistant |
| Andi | `andi_v3.1` | Male | American-born Chinese accent | voice assistant |
| An Yuqing | `anyuqing_v3.1` | Female | sweet young woman | voice assistant, narration, news reading |

## Premium English Voices

**15** voices, supporting English only.

| Name | voice | Gender | Accent |
| - | - | - | - |
| Emily | `Emily_v3.1` | Female | British |
| Luna | `Luna_v3.1` | Female | British |
| Eric | `Eric_v3.1` | Male | British |
| Luca | `Luca_v3.1` | Male | British |
| Abby | `Abby_v3.1` | Female | American |
| Annie | `Annie_v3.1` | Female | American |
| Ava | `Ava_v3.1` | Female | American |
| Beth | `Beth_v3.1` | Female | American |
| Betty | `Betty_v3.1` | Female | American |
| Cally | `Cally_v3.1` | Female | American |
| Cindy | `Cindy_v3.1` | Female | American |
| Donna | `Donna_v3.1` | Female | American |
| Andy | `Andy_v3.1` | Male | American |
| Brian | `Brian_v3.1` | Male | American |
| David | `David_v3.1` | Male | American |

## Other System Voices

**27** voices. Supported languages are not listed individually for this group; use text in your target language to check the synthesis quality.

| Name | voice | Gender | Voice qualities | Use cases |
| - | - | - | - | - |
| Long Anyuanfei | `longanyuanfei_v3.1` | Female | proud imperial consort | social companionship |
| Long Jielidou | `longjielidou_v3.1` | Male | innocent boy | children’s companionship |
| Long Anlingxi | `longanlingxi_v3.1` | Female | cute and sweet | social companionship |
| Long Huohuo | `longhuohuo_v3.1` | Male | mischievous teenage boy | character voice |
| Long Yingtao | `longyingtao_v3.1` | Female | gentle, calm woman | customer service |
| Long Anya | `longanya_v3.1` | Female | refined, elegant woman | social companionship |
| Long Wan | `longwan_v3.1` | Female | delicate, soft-spoken woman | social companionship |
| Long Xing | `longxing_v3.1` | Female | gentle girl next door | social companionship |
| Long Hua | `longhua_v3.1` | Female | energetic, sweet woman | social companionship |
| Long Han | `longhan_v3.1` | Male | warm, devoted man | social companionship |
| Long Anzhi | `longanzhi_v3.1` | Male | wise, moderately mature man | social companionship |
| Long Zhe | `longzhe_v3.1` | Male | awkward but warm man | social companionship |
| Long Anyang | `longanyang_v3.1` | Male | cheerful young man | social companionship |
| Li Bai | `libai_v3.1` | Male | classical poet | poetry recitation |
| Long Ling | `longling_v3.1` | Female | childlike, reserved girl | child voice |
| Long Niuniu | `longniuniu_v3.1` | Male | cheerful boy | children’s audiobooks |
| Long Shanshan | `longshanshan_v3.1` | Male | theatrical child voice | children’s audiobooks |
| Long Paopao | `longpaopao_v3.1` | Female | Powerpuff-style voice | children’s companionship |
| loongstella | `loongstella_v3.1` | Female | confident, brisk woman | news reading |
| Long Yuan | `longyuan_v3.1` | Female | warm, soothing woman | audiobooks |
| Long Miao | `longmiao_v3.1` | Female | expressive female intonation | audiobooks |
| Long Sanshu | `longsanshu_v3.1` | Male | composed, textured male voice | audiobooks |
| Long Anli | `longanli_v3.1` | Female | brisk, poised woman | voice assistant |
| Long Anwen | `longanwen_v3.1` | Female | elegant, articulate woman | voice assistant |
| Long Anlang | `longanlang_v3.1` | Male | clear, brisk man | voice assistant |
| Long Xiaoxia | `longxiaoxia_v3.1` | Female | composed, authoritative woman | voice assistant |
| Long Anchong | `longanchong_v3.1` | Male | enthusiastic salesman | live shopping |

## Custom Voices

For your own voice, use [Voice Enrollment](/docs/en/api-manual/audio-series/qwen-audio-tts/voice-enrollment): provide a recording to clone a voice, or a text description to design one. After the task completes, pass the full returned `voice` name to the speech synthesis API.

* Custom voices can only be used by the account that created them
* By default, they expire **6 hours** after the creation task completes; synthesis calls do not extend their lifetime
* After expiry, synthesis returns `404` (`voice_expired`); create a new voice with Voice Enrollment
* Voices created with the old `qwen-voice-design` model cannot be used here; recreate them with Voice Enrollment when migrating
