Skip to main content
Qwen-Audio-TTS

Qwen-Audio-TTS Python SDK

Use the DashScope Python SDK to integrate Qwen-Audio-TTS real-time speech synthesis into your application through non-streaming, one-way streaming, or bidirectional streaming modes.

Service endpoint

The SDK uses the Beijing region endpoint by default. To switch to a different region, modify dashscope.base_websocket_api_url before initialization.
  • Singapore
  • China (Beijing)
wss://{WorkspaceId}.ap-southeast-1.maas.aliyuncs.com/api-ws/v1/inferenceReplace {WorkspaceId} with your actual workspace ID.
Alibaba Cloud Model Studio has released workspace-specific domains for the China (Beijing) and Singapore regions. The new dedicated domains deliver superior performance and higher stability for inference requests. We recommend migrating to the new domains:
  • China (Beijing): from dashscope.aliyuncs.com to {WorkspaceId}.cn-beijing.maas.aliyuncs.com
  • Singapore: from dashscope-intl.aliyuncs.com to {WorkspaceId}.ap-southeast-1.maas.aliyuncs.com

SpeechSynthesizer

Package path: dashscope.audio.tts_v2.SpeechSynthesizer

Constructor

SpeechSynthesizer(
    model: str,
    voice: str,
    format: AudioFormat = AudioFormat.MP3_22050HZ_MONO_256KBPS,
    volume: int = 50,
    speech_rate: float = 1.0,
    pitch_rate: float = 1.0,
    callback: ResultCallback = None)

call() - non-streaming

Method signature:
def call(self, text: str) -> bytes
Parameters:
ParameterTypeRequiredDescription
textstrYesThe full text to synthesize. Maximum length: 20,000 characters.
Return value: bytes containing the complete audio data. Description: This blocking call returns the complete audio data at once. It is best suited for short text where real-time streaming is not required. Reinitialize the SpeechSynthesizer instance before each call.

streaming_call() - streaming

Method signature:
def streaming_call(self, text: str) -> None
Parameters:
ParameterTypeRequiredDescription
textstrYesA text segment to synthesize. Call this method multiple times to append text. Maximum per call: 20,000 characters. Cumulative maximum: 200,000 characters.
Description: This bidirectional streaming call accepts text in segments and delivers synthesized audio through callbacks in real time. It is ideal for integration with large language models where text is generated progressively. Call streaming_complete() after sending all text.

streaming_complete() - end streaming

Method signature:
def streaming_complete(self) -> None
Description: Notifies the server that all text has been sent. Blocks the current thread until the remaining text is synthesized and all audio data is returned. Failing to call this method may result in trailing text not being converted to speech.

streaming_cancel() - cancel streaming synthesis

Method signature:
def streaming_cancel(self, complete_timeout_millis: int = 10000) -> None
Parameters:
ParameterTypeRequiredDescription
complete_timeout_millisintNoTimeout in milliseconds for waiting for the server to return the task-finished event. Default value: 10000.
Description: Cancels the current streaming speech synthesis turn. After calling this method, the SDK immediately ends the current task. You can start a new synthesis task on the same connection without reinitializing the SpeechSynthesizer instance.
Version requirement: This feature requires Python SDK 1.26.4 or later.

get_last_request_id() - get request ID

Method signature:
def get_last_request_id(self) -> str
Return value: str containing the request ID of the most recent request. Use this for troubleshooting and tracing.

get_first_package_delay() - get first-packet latency

Method signature:
def get_first_package_delay(self) -> int
Return value: int representing the delay in milliseconds from sending text to receiving the first audio chunk. Call this after synthesis completes.

get_response() - get response message

Method signature:
def get_response(self) -> str
Return value: str containing the JSON-formatted response message from the most recent synthesis task, including request status and output information.

Constructor parameters

The following parameters are set through the SpeechSynthesizer constructor to control the model, voice, format, and audio characteristics.
ParameterTypeRequiredDescription
modelstrYesThe model name.
voicestrYesThe voice used for speech synthesis.
  • System voices: See Qwen-Audio-TTS voice list
  • Cloned voices: Custom voices created through voice cloning
  • Custom voices: Custom voices created through voice design
formatenumNoAudio encoding format and sample rate.Default: AudioFormat.MP3_22050HZ_MONO_256KBPS.The AudioFormat enum is located in dashscope.audio.tts_v2 and supports MP3, WAV, PCM, and other formats.
volumeintNoThe volume level.Default value: 50.Valid values: [0, 100].
speech_ratefloatNoThe speech rate.Default value: 1.0.Valid values: [0.5, 2.0].
pitch_ratefloatNoThe pitch.Default value: 1.0.Valid values: [0.5, 2.0].
bit_rateintNoThe audio bit rate in kbps. When the audio format is mp3 or opus, use bit_rate to adjust the bit rate.Default value: 32.Valid values: [6, 510].Set bit_rate through the additional_params parameter:
synthesizer = SpeechSynthesizer(
          model="qwen-audio-3.0-tts-flash",
          voice="longanhuan_v3.6",
          additional_params={"bit_rate": 128}
      )
word_timestamp_enabledboolNoSpecifies whether to enable word-level timestamps.Default value: false.Available only in streaming output mode. Cloned voices are supported. For supported system voices, see Qwen-Audio-TTS voice list.Set word_timestamp_enabled through the additional_params parameter:
synthesizer = SpeechSynthesizer(
          model="qwen-audio-3.0-tts-flash",
          voice="your_voice",  # A system or cloned voice that supports word-level timestamps
          additional_params={"word_timestamp_enabled": True}
      )
seedintNoA random seed for controlling variation in the synthesis output. When the model version, text, voice, and other parameters are unchanged, using the same seed produces identical results.Default value: 0.Valid values: [0, 65535].
language_hintslist[str]No
  • This parameter is an array, but the current version only processes the first element. Pass a single value.
  • This parameter specifies the target language for speech synthesis. It's unrelated to the language of the audio sample used in voice cloning. To set the source language for a cloning task, see the voice cloning API reference.
Specifies the target language for speech synthesis to improve output quality.When digit pronunciation, abbreviation expansion, symbol reading, or minority-language synthesis doesn't meet expectations, use this parameter. For example:
  • Unexpected digit pronunciation: "hello, this is 110" is read as "hello, this is one zero" instead of the expected Chinese pronunciation
  • Inaccurate symbol pronunciation: "@" is read as the Chinese equivalent instead of "at"
  • Poor minor language synthesis quality with unnatural results
  • zh: Chinese
  • en: English
  • fr: French
  • de: German
  • ja: Japanese
  • ko: Korean
  • ru: Russian
  • pt: Portuguese
  • th: Thai
  • id: Indonesian
  • vi: Vietnamese
  • es: Spanish
  • it: Italian
  • ms: Malaysian
  • fil: Filipino
  • ar: Arabic
instructionstrNoControls synthesis characteristics such as dialect, emotion, or speaking style.For usage details, see Instruction control.
enable_aigc_tagboolNoSpecifies whether to embed an AIGC watermark in the generated audio. When set to true, the watermark is embedded in audio files of supported formats (wav/mp3/opus).Default value: false.Set enable_aigc_tag, aigc_propagator, and aigc_propagate_id through the additional_params parameter:
synthesizer = SpeechSynthesizer(
          model="qwen-audio-3.0-tts-flash",
          voice="longanhuan_v3.6",
          additional_params={
              "enable_aigc_tag": True,
              "aigc_propagator": "your_propagator",
              "aigc_propagate_id": "your_propagate_id"
          }
      )
aigc_propagatorstrNoSets the ContentPropagator field in the AIGC watermark, identifying the content propagator. Takes effect only when enable_aigc_tag is true.Default value: Alibaba Cloud UID.Set through the additional_params parameter. See the enable_aigc_tag example.
aigc_propagate_idstrNoSets the PropagateID field in the AIGC watermark, uniquely identifying a specific propagation action. Takes effect only when enable_aigc_tag is true.Default value: The request ID of the current speech synthesis request.Set through the additional_params parameter. See the enable_aigc_tag example.
hot_fixdictNoConfigures pronunciation corrections and text replacements applied before synthesis.Parameters:
  • pronunciation: Custom pronunciation. Specifies pinyin annotations for words to correct inaccurate default pronunciations.
  • replace: Text replacement. Replaces specified words with target text before synthesis. The replaced text is used as the actual synthesis input.
Example:
synthesizer = SpeechSynthesizer(
    model="qwen-audio-3.0-tts-flash",
    voice="your_voice_id", # Voice
    hot_fix={
        "pronunciation": [{"weather": "tian1 qi4"}],
        "replace": [{"today": "gold day"}]
    }
)
callbackResultCallbackNoA callback instance for receiving synthesized audio and event notifications asynchronously. When set, call() runs in streaming mode and delivers audio through the on_data callback. When not set, call() runs in non-streaming mode and returns the complete audio as bytes.

ResultCallback

Package path: dashscope.audio.tts_v2.ResultCallback

on_open() - connection established

Method signature:
def on_open(self) -> None
Triggered when: The WebSocket connection is successfully established. Use this callback to initialize audio output streams or open file resources.

on_event() - receive server response

Method signature:
def on_event(self, message: str) -> None
Parameters:
ParameterTypeRequiredDescription
messagestrYesA server response event in JSON format containing header (request information) and payload (output information). The payload.output field contains event type, original text, and other details. See output field in on_event messages.
Triggered when: A server response is received. The message is a JSON string containing synthesis event output (event type, original text, sentence information). Parse with json.loads(message) and access payload.output for details.

on_complete() - synthesis complete

Method signature:
def on_complete(self) -> None
Triggered when: All text has been synthesized and all audio data has been delivered through on_data. Use this callback to call get_first_package_delay() for performance metrics.

on_data() - receive audio data

Method signature:
def on_data(self, data: bytes) -> None
Parameters:
ParameterTypeRequiredDescription
databytesYesA chunk of audio binary data in the format specified by the constructor's format parameter.
Triggered when: An audio data chunk is received. This callback is invoked multiple times during synthesis. Use it to write data to a file or feed it to a playback device.

on_error() - error occurred

Method signature:
def on_error(self, message: str) -> None
Parameters:
ParameterTypeRequiredDescription
messagestrYesAn error description containing the error code and detailed reason.
Triggered when: An error occurs during synthesis. The connection closes automatically after this callback fires. Log the error for troubleshooting.

on_close() - connection closed

Method signature:
def on_close(self) -> None
Triggered when: The WebSocket connection closes (whether normally or due to an error). Use this callback to release resources such as audio playback devices.

output field in on_event messages

The JSON message received by the on_event callback contains a payload.output field with synthesis event information. Use this field to track synthesis progress and retrieve per-sentence details. The output field structure is as follows:
FieldTypeDescription
typestrEvent type. Values: sentence-begin (sentence synthesis started), sentence-synthesis (sentence synthesis in progress), or sentence-end (sentence synthesis finished).
original_textstrThe original text of the current sentence. Returned in sentence-begin and sentence-end events.
sentencedictSentence information. Contains index (sentence sequence number) and words (word list with timestamp information when word_timestamp_enabled is on).
Message example:
{
  "header": {
    "task_id": "xxx",
    "event": "result-generated",
    "attributes": {}
  },
  "payload": {
    "output": {
      "type": "sentence-begin",
      "original_text": "How is the weather today?",
      "sentence": {
        "index": 0,
        "words": []
      }
    }
  }
}
Parsing example:
import json

def on_event(self, message):
    data = json.loads(message)
    output = data.get('payload', {}).get('output', {})
    event_type = output.get('type', '')
    original_text = output.get('original_text', '')
    if event_type:
        print(f'Event type: {event_type}, Original text: {original_text}')

Code examples

The SDK supports the following synthesis modes:
  • Non-streaming: A blocking call that sends the complete text at once and returns the full audio directly. Best suited for short-text speech synthesis.
  • Unidirectional streaming: A non-blocking call that sends the complete text at once and delivers audio data (potentially in chunks) through a callback function. Best suited for short-text scenarios that require low latency.
  • Bidirectional streaming: A non-blocking call that sends text in multiple segments and delivers incrementally synthesized audio through a callback function in real time. Best suited for long-text scenarios that require low latency.
  • Non-streaming
  • One-way streaming
  • Bidirectional streaming
The text sent in a single call must not exceed 20,000 characters. Exceeding this limit causes an error.
Reinitialize the SpeechSynthesizer instance before each call.
# coding=utf-8

import dashscope
from dashscope.audio.tts_v2 import *
import os

# The API Keys for Singapore and Beijing regions are different. Get API Key: https://www.alibabacloud.com/help/model-studio/get-api-key
# If environment variable is not configured, replace the following line with your Model Studio API Key: dashscope.api_key = "sk-xxx"
dashscope.api_key = os.environ.get('DASHSCOPE_API_KEY')

# The following configuration is for the Singapore region. Replace "{WorkspaceId}" with your actual workspace ID. The configuration varies by region.
dashscope.base_websocket_api_url='wss://{WorkspaceId}.ap-southeast-1.maas.aliyuncs.com/api-ws/v1/inference'

# Model
model = "qwen-audio-3.0-tts-flash"
# Voice
voice = "longanhuan_v3.6"

# Instantiate SpeechSynthesizer, passing request parameters such as model and voice in the constructor
synthesizer = SpeechSynthesizer(model=model, voice=voice)
# Send text for synthesis and get binary audio
audio = synthesizer.call("How is the weather today?")
# The first text submission requires establishing a WebSocket connection, so the first packet latency includes the connection setup time
print('[Metric] requestId: {}, first packet latency: {} ms'.format(
    synthesizer.get_last_request_id(),
    synthesizer.get_first_package_delay()))

# Save audio to local file
with open('output.mp3', 'wb') as f:
    f.write(audio)
  • dashscope CLI
Call via the dashscope command line.
export DASHSCOPE_API_KEY="your-api-key"
# Replace {WorkspaceId} with your workspace ID, ap-southeast-1 with your region (e.g., us-east-1, eu-central-1)
export DASHSCOPE_HTTP_BASE_URL="https://{WorkspaceId}.ap-southeast-1.maas.aliyuncs.com/api/v1"
dashscope speech-synthesis create -m qwen-audio-3.0-tts-flash -t "Hello world" --voice longanhuan_v3.6
The SDK Expert interactive assistant can accomplish the same development and troubleshooting via natural language, see DashScope SDK Expert.
For the complete region table, see Base URL overview.
Text Generation
Image Generation
  • FAQ
Video Generation
World models
Audio
  • Audio generation
Realtime API
Text Embedding
Decision Model
TokenPlan
Model Production