model | str | Yes | The model name. |
voice | str | Yes | The voice used for speech synthesis.
- System voices: See Qwen-Audio-TTS voice list
- Cloned voices: Custom voices created through voice cloning
- Custom voices: Custom voices created through voice design
|
format | enum | No | Audio encoding format and sample rate.Default: AudioFormat.MP3_22050HZ_MONO_256KBPS.The AudioFormat enum is located in dashscope.audio.tts_v2 and supports MP3, WAV, PCM, and other formats. |
volume | int | No | The volume level.Default value: 50.Valid values: [0, 100]. |
speech_rate | float | No | The speech rate.Default value: 1.0.Valid values: [0.5, 2.0]. |
pitch_rate | float | No | The pitch.Default value: 1.0.Valid values: [0.5, 2.0]. |
bit_rate | int | No | The audio bit rate in kbps. When the audio format is mp3 or opus, use bit_rate to adjust the bit rate.Default value: 32.Valid values: [6, 510].Set bit_rate through the additional_params parameter:synthesizer = SpeechSynthesizer(
model="qwen-audio-3.0-tts-flash",
voice="longanhuan_v3.6",
additional_params={"bit_rate": 128}
)
|
word_timestamp_enabled | bool | No | Specifies whether to enable word-level timestamps.Default value: false.Available only in streaming output mode. Cloned voices are supported. For supported system voices, see Qwen-Audio-TTS voice list.Set word_timestamp_enabled through the additional_params parameter:synthesizer = SpeechSynthesizer(
model="qwen-audio-3.0-tts-flash",
voice="your_voice", # A system or cloned voice that supports word-level timestamps
additional_params={"word_timestamp_enabled": True}
)
|
seed | int | No | A random seed for controlling variation in the synthesis output. When the model version, text, voice, and other parameters are unchanged, using the same seed produces identical results.Default value: 0.Valid values: [0, 65535]. |
language_hints | list[str] | No |
- This parameter is an array, but the current version only processes the first element. Pass a single value.
- This parameter specifies the target language for speech synthesis. It's unrelated to the language of the audio sample used in voice cloning. To set the source language for a cloning task, see the voice cloning API reference.
Specifies the target language for speech synthesis to improve output quality.When digit pronunciation, abbreviation expansion, symbol reading, or minority-language synthesis doesn't meet expectations, use this parameter. For example:
- Unexpected digit pronunciation: "hello, this is 110" is read as "hello, this is one zero" instead of the expected Chinese pronunciation
- Inaccurate symbol pronunciation: "@" is read as the Chinese equivalent instead of "at"
- Poor minor language synthesis quality with unnatural results
- zh: Chinese
- en: English
- fr: French
- de: German
- ja: Japanese
- ko: Korean
- ru: Russian
- pt: Portuguese
- th: Thai
- id: Indonesian
- vi: Vietnamese
- es: Spanish
- it: Italian
- ms: Malaysian
- fil: Filipino
- ar: Arabic
|
instruction | str | No | Controls synthesis characteristics such as dialect, emotion, or speaking style.For usage details, see Instruction control. |
enable_aigc_tag | bool | No | Specifies whether to embed an AIGC watermark in the generated audio. When set to true, the watermark is embedded in audio files of supported formats (wav/mp3/opus).Default value: false.Set enable_aigc_tag, aigc_propagator, and aigc_propagate_id through the additional_params parameter:synthesizer = SpeechSynthesizer(
model="qwen-audio-3.0-tts-flash",
voice="longanhuan_v3.6",
additional_params={
"enable_aigc_tag": True,
"aigc_propagator": "your_propagator",
"aigc_propagate_id": "your_propagate_id"
}
)
|
aigc_propagator | str | No | Sets the ContentPropagator field in the AIGC watermark, identifying the content propagator. Takes effect only when enable_aigc_tag is true.Default value: Alibaba Cloud UID.Set through the additional_params parameter. See the enable_aigc_tag example. |
aigc_propagate_id | str | No | Sets the PropagateID field in the AIGC watermark, uniquely identifying a specific propagation action. Takes effect only when enable_aigc_tag is true.Default value: The request ID of the current speech synthesis request.Set through the additional_params parameter. See the enable_aigc_tag example. |
hot_fix | dict | No | Configures pronunciation corrections and text replacements applied before synthesis.Parameters:
- pronunciation: Custom pronunciation. Specifies pinyin annotations for words to correct inaccurate default pronunciations.
- replace: Text replacement. Replaces specified words with target text before synthesis. The replaced text is used as the actual synthesis input.
Example:synthesizer = SpeechSynthesizer(
model="qwen-audio-3.0-tts-flash",
voice="your_voice_id", # Voice
hot_fix={
"pronunciation": [{"weather": "tian1 qi4"}],
"replace": [{"today": "gold day"}]
}
)
|
callback | ResultCallback | No | A callback instance for receiving synthesized audio and event notifications asynchronously. When set, call() runs in streaming mode and delivers audio through the on_data callback. When not set, call() runs in non-streaming mode and returns the complete audio as bytes. |