Generate speech, sound effects, and ambient sound together from text descriptions and reference audio to create complete audio scenes. Explore examples and prompt techniques for dialogue and podcasts.
Use cases
Audio generation goes beyond converting text to speech. It can generate speech, sound effects, and ambient sound together to create complete audio scenes:
- Speech and dialogue: narration, multi-speaker dialogue, podcasts, audiobooks, and audio dramas.
- Audio scenes: dialogue accompanied by rain, ocean waves, or game sound effects.
Examples
Explore prompts and generated audio for different scenarios. Results may vary between requests with the same prompt.
| Scenario | Prompt | Audio |
|---|---|---|
| Game voice-overs | ||
| Film dialogue |
Write a prompt
Describe the required sounds in one prompt, distinguishing dialogue from scene directions.
| Element | Example |
|---|---|
| Task | “An opening exchange between two podcast hosts” or “An audio drama scene with dialogue in the rain” |
| Characters and voices | “Character A has a low male voice; character B has a bright female voice” |
| Dialogue | Put spoken lines in quotation marks and identify the speaker |
| Delivery | “Whisper”, “sound surprised”, or “pause briefly before continuing” |
| Background and effects | “Rain outside the window with quiet page-turning sounds indoors” |
| Sound transitions | “Footsteps approach before the character starts speaking” |
Use reference audio
Provide publicly accessible audio URLs or Base64-encoded audio. Use @voice1, @voice2, and @voice3 in the prompt to refer to clips in the order provided.
For example, with two reference clips:
qwen-audio-3.1-tts-next accepts up to three clips, each no longer than 30 seconds and no larger than 10 MB. Supported reference formats are WAV, MP3, and OGG Opus; raw PCM is not supported. For each clip, provide either a URL or Base64 data. The service must be able to access any URL you provide.
Supported models
| Model | Languages | Maximum input length | Maximum output duration per request |
|---|---|---|---|
| qwen-audio-3.1-tts-next | Chinese and English | 3,000 characters | Podcasts: 240 seconds (4 minutes); other scenarios: 120 seconds |
Integration and billing
- For authentication, parameters, and reference audio requests, see Audio generation API reference.
- Billing uses the original generated audio duration. Changing the speech rate does not change the billable duration. See Model pricing.