run-task
Starts a speech synthesis task and configures the model, voice, sample rate, and other parameters.
When to send: Immediately after the WebSocket connection is established.
Response event: The server returns a task-started event. Wait for this event before sending subsequent commands.
header object(required)
Properties action string(required)The command type. Set to run-task.task_id string(required)A client-generated task ID in UUID format. This ID correlates subsequent events and must match the task_id in the continue-task and finish-task commands.streaming string(required)Set to duplex.object(required)
Properties task_group string(required)The task group. Set to audio.task string(required)The task type. Set to tts.function string(required)The function type. Set to SpeechSynthesizer.model string(required)The model name.input object(required)Set to an empty object {}. Send the text to synthesize through the continue-task command.parameters object(required)Speech synthesis parameters.
Properties text_type string(required)Set to PlainText.voice string(required)The voice used for speech synthesis.
string(optional)The audio encoding format.Valid values:
integer(optional)The audio sample rate in Hz.Valid values: 8000, 16000, 22050 (default), 24000, 44100, 48000.volume integer(optional)The volume level.Default value: 50.Valid values: [0, 100].rate float(optional)The speech rate.Default value: 1.0.Valid values: [0.5, 2.0].pitch float(optional)The pitch.Default value: 1.0.Valid values: [0.5, 2.0].bit_rate integer(optional)The audio bit rate in kbps. When the audio format is mp3 or opus, use bit_rate to adjust the bit rate.Default value: 32.Valid values: [6, 510].enable_ssml boolean(optional)Specifies whether to enable SSML.Default value: false.When set to true, only one continue-task command is allowed.For the SSML usage restrictions (supported models, voices, and APIs), see Limitations.word_timestamp_enabled boolean(optional)Specifies whether to enable word-level timestamps.Default value: false.Available only in streaming output mode. Cloned voices are supported. For supported system voices, see Qwen-Audio-TTS voice list.seed integer(optional)A random seed for controlling variation in the synthesis output. When the model version, text, voice, and other parameters are unchanged, using the same seed produces identical results.Default value: 0.Valid values: [0, 65535].language_hints array[string](optional)Specifies the target language for speech synthesis to improve output quality.When digit pronunciation, abbreviation expansion, symbol reading, or minority-language synthesis doesn't meet expectations, use this parameter. For example:
Valid values
string(optional)Sets an instruction to control dialect, emotion, or voice character during synthesis. For detailed usage, see Instruction control.enable_aigc_tag boolean(optional)Specifies whether to embed an AIGC watermark in the generated audio. When set to true, the watermark is embedded in audio files of supported formats (wav/mp3/opus).Default value: false.aigc_propagator string(optional)Sets the ContentPropagator field in the AIGC watermark, identifying the content propagator. Takes effect only when enable_aigc_tag is true.Default value: Alibaba Cloud UID.aigc_propagate_id string(optional)Sets the PropagateID field in the AIGC watermark, uniquely identifying a specific propagation action. Takes effect only when enable_aigc_tag is true.Default value: The request ID of the current speech synthesis request.hot_fix object(optional)Configures pronunciation corrections and text replacements applied before synthesis.Parameters:
|
continue-task
Sends the text to synthesize. The text can be sent all at once or in multiple segments.
When to send: After receiving the task-started event from the server.
Limits:
- Maximum of 20,000 characters per message
- Maximum of 200,000 characters cumulatively
- The send interval must not exceed 23 seconds; otherwise, the connection times out.
header object(required)
Properties action string(required)The command type. Set to continue-task.task_id string(required)The task ID in UUID format. Must match the task_id in run-task.streaming string(required)Set to duplex.object(required)
Properties input object(required)Contains the text to synthesize.text string(required)The text to synthesize. Maximum of 20,000 characters per message and 200,000 characters cumulatively. |
finish-task
Notifies the server that all text has been sent and requests task completion.
When to send: Immediately after all text has been sent.
Response event: The server returns a task-finished event.
header object(required)
Properties action string(required)The command type. Set to finish-task.task_id string(required)The task ID in UUID format. Must match the task_id in run-task.streaming string(required)Set to duplex.object(required)
Properties input object(required)Set to {} for normal task completion. Include directive to cancel the current synthesis turn.directive string(optional)Controls how the task ends. Currently only cancel is supported. When set to cancel, the current synthesis turn is canceled and the server immediately returns a task-finished event without producing further audio.After cancellation, you can start a new synthesis task on the same WebSocket connection by sending a new run-task event without reconnecting. |