Synthesize speech with Qwen-Audio-TTS using the DashScope Java SDK.
Service endpoints
By default, the SDK connects to the Beijing region endpoint. To use a different region, set Constants.baseWebsocketApiUrl before initializing the SDK.
- Singapore
- China (Beijing)
wss://{WorkspaceId}.ap-southeast-1.maas.aliyuncs.com/api-ws/v1/inferenceReplace {WorkspaceId} with your actual workspace ID.SpeechSynthesizer
Package: com.alibaba.dashscope.audio.ttsv2.SpeechSynthesizer
Constructor
param: Speech synthesis parameters, built withSpeechSynthesisParam.builder()callback: Callback for streaming calls. Pass null for non-streaming calls.
call() - Non-streaming/unidirectional streaming synthesis
Method signature:
| Parameter | Type | Required | Description |
|---|---|---|---|
text | String | Yes | Text to synthesize. Maximum length: 20,000 characters. |
ByteBuffer or null. For non-streaming calls, returns the complete audio data. For unidirectional streaming calls, this method returns null; audio is delivered through the callback.
streamingCall() - Bidirectional streaming synthesis
Method signature:
| Parameter | Type | Required | Description |
|---|---|---|---|
text | String | Yes | Text to synthesize. Maximum length: 20,000 characters. You can call this method multiple times to append text. |
streamingComplete() - End bidirectional streaming
Method signature:
streamingCancel() - Cancel bidirectional streaming
Method signature:
SpeechSynthesizer instance.
callAsFlowable() - Unidirectional streaming synthesis (reactive)
Method signature:
| Parameter | Type | Required | Description |
|---|---|---|---|
text | String | Yes | Text to synthesize. |
streamingCallAsFlowable() - Bidirectional streaming synthesis (reactive)
Method signature:
| Parameter | Type | Required | Description |
|---|---|---|---|
textStream | Flowable<String> | Yes | Reactive stream of text. |
SpeechSynthesisResult> reactive stream.
getDuplexApi().close() - Close WebSocket connection
Method signature:
| Parameter | Type | Required | Description |
|---|---|---|---|
code | int | Yes | Close code. |
reason | String | Yes | Close reason. |
boolean. Returns true if the connection was closed successfully, false otherwise.
getLastRequestId() - Get request ID
Method signature:
String, the request ID.
getFirstPackageDelay() - Get first-packet latency
Method signature:
long. The first-packet latency in milliseconds, measured from sending the first text segment to receiving the first audio packet.
SpeechSynthesisParam
Package: com.alibaba.dashscope.audio.ttsv2.SpeechSynthesisParam
Example:
Builder methods
| Method | Parameter type | Required | Description |
|---|---|---|---|
model(String) | String | Yes | The model name. |
voice(String) | String | Yes | The voice used for speech synthesis.
|
format(SpeechSynthesisAudioFormat) | enum | No | Audio encoding format and sample rate.Default: SpeechSynthesisAudioFormat.MP3_22050HZ_MONO_256KBPS.SpeechSynthesisAudioFormat package: com.alibaba.dashscope.audio.ttsv2.SpeechSynthesisAudioFormat. |
volume(int) | int | No | The volume level.Default value: 50.Valid values: [0, 100]. |
speechRate(float) | float | No | The speech rate.Default value: 1.0.Valid values: [0.5, 2.0]. |
pitchRate(float) | float | No | The pitch.Default value: 1.0.Valid values: [0.5, 2.0]. |
enableWordTimestamp(boolean) | boolean | No | Specifies whether to enable word-level timestamps.Default value: false.Available only in streaming output mode. Cloned voices are supported. For supported system voices, see Qwen-Audio-TTS voice list. |
seed(int) | int | No | A random seed for controlling variation in the synthesis output. When the model version, text, voice, and other parameters are unchanged, using the same seed produces identical results.Default value: 0.Valid values: [0, 65535].For SDK versions earlier than 2.21.7, set seed through additional parameters. |
languageHints(List<String>) | List<String> | No | Specifies the target language for speech synthesis to improve output quality.When digit pronunciation, abbreviation expansion, symbol reading, or minority-language synthesis doesn't meet expectations, use this parameter. For example:
Valid values
|
instruction(String) | String | No | Controls synthesis characteristics such as dialect, emotion, or speaking style.For usage details, see Instruction control. |
hotFix(ParamHotFix) | ParamHotFix | No | Configures pronunciation corrections and text replacements applied before synthesis.Parameters:
|
parameter(String key, Object value) | String, Object | No | Sets Additional parameters. |
parameters(Map<String, Object>) | Map | No | Sets Additional parameters. |
Additional parameters
Set through parameter() or parameters().
Example:
| Parameter | Type | Required | Description |
|---|---|---|---|
bit_rate | integer | No | The audio bit rate in kbps. When the audio format is mp3 or opus, use bit_rate to adjust the bit rate.Default value: 32.Valid values: [6, 510]. |
enable_aigc_tag | boolean | No | Specifies whether to embed an AIGC watermark in the generated audio. When set to true, the watermark is embedded in audio files of supported formats (wav/mp3/opus).Default value: false. |
aigc_propagator | String | No | Sets the ContentPropagator field in the AIGC watermark, identifying the content propagator. Takes effect only when enable_aigc_tag is true.Default value: Alibaba Cloud UID. |
aigc_propagate_id | String | No | Sets the PropagateID field in the AIGC watermark, uniquely identifying a specific propagation action. Takes effect only when enable_aigc_tag is true.Default value: The request ID of the current speech synthesis request. |
ResultCallback
Package: com.alibaba.dashscope.common.ResultCallback
onEvent() - Receive audio data
Method signature:
| Parameter | Type | Required | Description |
|---|---|---|---|
result | SpeechSynthesisResult | Yes | Triggered when a synthesis event is received. Contains the audio frame, timestamp information, and output information (event type, original text, etc.). |
onComplete() - Synthesis complete
Method signature:
onError() - Error handling
Method signature:
| Parameter | Type | Required | Description |
|---|---|---|---|
e | Exception | Yes | Triggered when an error occurs. Contains the exception information. |
SpeechSynthesisResult
Package: com.alibaba.dashscope.audio.tts.SpeechSynthesisResult
getAudioFrame() - Get audio data frame
Method signature:
ByteBuffer, the audio data frame.
getTimestamp() - Get timestamp information
Method signature:
getOutput() - Get output information
Method signature:
com.google.gson.JsonObject, the output information of the synthesis event, containing event type and text content. Requires SDK version >= 2.22.0.
Sentence-level timestamp information (Sentence)
Sentence encapsulates sentence-level timestamp information.
getBeginTime() - Get sentence start time
Method signature:
getEndTime() - Get sentence end time
Method signature:
getWords() - Get word-level timestamps
Method signature:
List of Word objects containing word-level timestamp information. May be empty.
Word-level timestamp information (Word)
Word encapsulates word-level timestamp information.
getBeginTime() - Get word start time
Method signature:
getEndTime() - Get word end time
Method signature:
getText() - Get text
Method signature:
String, the text content.
getPhonemes() - Get phoneme-level timestamps
Method signature:
List of Phoneme objects containing phoneme-level timestamp information. May be empty.
Phoneme-level timestamp information (Phoneme)
Phoneme encapsulates phoneme-level timestamp information.
getBeginTime() - Get phoneme start time
Method signature:
getEndTime() - Get phoneme end time
Method signature:
getText() - Get text
Method signature:
String, the text content.
getTone() - Get tone
Method signature:
- In English, 0, 1, and 2 represent unstressed, primary stress, and secondary stress respectively.
- In Chinese pinyin, 1, 2, 3, 4, and 5 represent the first, second, third, fourth, and neutral tones respectively.
Output information (output)
getOutput() returns a JsonObject that encapsulates synthesis event output information. Retrieve it in the onEvent callback or the Flowable stream. It contains the following fields:
| Field | Type | Description |
|---|---|---|
type | String | Event type. Possible values: sentence-begin (sentence start; returns the text to be synthesized), sentence-synthesis (audio synthesis in progress; returns an audio data chunk), sentence-end (sentence end; returns text content and word-level timestamps). |
original_text | String | Original text of the current sentence. Returned in sentence-begin and sentence-end events. |
sentence | JsonObject | Sentence information, containing the sentence index (index) and word-level timestamps (words). The sentence-end event includes full word-level timestamp information. |
Sample code
The SDK supports the following synthesis modes:
- Non-streaming: A blocking call that sends the complete text at once and returns the full audio directly. Best suited for short-text speech synthesis.
- Unidirectional streaming: A non-blocking call that sends the complete text at once and delivers audio data (potentially in chunks) through a callback function. Best suited for short-text scenarios that require low latency.
- Bidirectional streaming: A non-blocking call that sends text in multiple segments and delivers incrementally synthesized audio through a callback function in real time. Best suited for long-text scenarios that require low latency.
- Non-streaming calls
- Unidirectional streaming calls
- Bidirectional streaming calls
Flowable-based calls
Flowable is an RxJava type representing a reactive stream that supports backpressure. For more information, see RxJava Flowable documentation.
Before using Flowable, make sure the RxJava library is integrated and you understand the basics of reactive programming.
The text length per individual call must not exceed 20,000 characters, and the cumulative text length across all calls must not exceed 200,000 characters.
- Unidirectional streaming calls
- Bidirectional streaming calls
The following example shows how to use the
blockingForEach interface of a Flowable object to retrieve each streamed SpeechSynthesisResult object in a blocking manner.