Skip to main content
Model Playground

Audio generation

Generate speech, sound effects, and ambient sound together from text descriptions and reference audio to create complete audio scenes. Explore examples and prompt techniques for dialogue and podcasts.

Use cases

Audio generation goes beyond converting text to speech. It can generate speech, sound effects, and ambient sound together to create complete audio scenes:
  • Speech and dialogue: narration, multi-speaker dialogue, podcasts, audiobooks, and audio dramas.
  • Audio scenes: dialogue accompanied by rain, ocean waves, or game sound effects.
For converting a fixed script into speech, you can also use Text-to-speech. Use audio generation when you need to describe both speakers and other sounds in a scene.

Examples

Explore prompts and generated audio for different scenarios. Results may vary between requests with the same prompt.
ScenarioPromptAudio
Game voice-overs
NPC dialogue for a battle scene in a game. At a temporary aid station on the edge of a battlefield, a medic treats wounded soldiers and calls out over the chaos. Background sounds include distant battle cries, wounded soldiers groaning, and fabric tearing. The medic is a thirty-five-year-old woman with an urgent, resolute voice, full of anxiety and concern, speaking quickly. She shouts in English: "Quick! Get the wounded down here! There's another one! Stay with me. Don't close your eyes! A stretcher! Where's the stretcher? Keep pressure on the wound! A medic is on the way! All of you, hang in there!"
Film dialogue
An English-language film scene on an almost empty railway platform at night. Rain ticks against a metal roof, and an idling train rumbles in the distance. Owen, a man around thirty with a tired, low voice, speaks to his older sister Ruth, whose voice is firm but trembling underneath. Owen says quietly, "You didn't have to come." Ruth replies, "You left your scarf." A brief rustle of fabric. Owen gives a faint, breathy laugh. "It's summer." Ruth says, "It gets cold where you're going." A departure whistle sounds far away. Owen's voice softens: "I'll call when I get there." Ruth takes a breath and answers, "Call before that." Footsteps move toward the train; keep the rain audible as the dialogue ends. No narrator and no background score.

Write a prompt

Describe the required sounds in one prompt, distinguishing dialogue from scene directions.
ElementExample
Task“An opening exchange between two podcast hosts” or “An audio drama scene with dialogue in the rain”
Characters and voices“Character A has a low male voice; character B has a bright female voice”
DialoguePut spoken lines in quotation marks and identify the speaker
Delivery“Whisper”, “sound surprised”, or “pause briefly before continuing”
Background and effects“Rain outside the window with quiet page-turning sounds indoors”
Sound transitions“Footsteps approach before the character starts speaking”
Start with the core dialogue and sounds, then add necessary style instructions. Use consistent character names and avoid conflicting voice or emotion instructions. Use sequencing instructions to describe transitions between sounds.

Use reference audio

Provide publicly accessible audio URLs or Base64-encoded audio. Use @voice1, @voice2, and @voice3 in the prompt to refer to clips in the order provided. For example, with two reference clips:
@voice1 cheerfully says: "What a lovely day." @voice2 laughs nearby.
qwen-audio-3.1-tts-next accepts up to three clips, each no longer than 30 seconds and no larger than 10 MB. Supported reference formats are WAV, MP3, and OGG Opus; raw PCM is not supported. For each clip, provide either a URL or Base64 data. The service must be able to access any URL you provide.

Supported models

ModelLanguagesMaximum input lengthMaximum output duration per request
qwen-audio-3.1-tts-nextChinese and English3,000 charactersPodcasts: 240 seconds (4 minutes); other scenarios: 120 seconds
For pricing and limits, see the model details.

Integration and billing

  • For authentication, parameters, and reference audio requests, see Audio generation API reference.
  • Billing uses the original generated audio duration. Changing the speech rate does not change the billable duration. See Model pricing.
Token Plan
Model Playground
  • Audio generation
  • Music generation
Statistics and Monitoring
Support