Skip to main content
Audio Generation

qwen-audio-3.1-tts-next

Qwen-Audio-3.1-TTS-Next is an AudioGen model designed for unified audio generation. Going beyond conventional TTS systems that focus solely on speech synthesis, it can generate complete audio content in a single pass from inputs such as text, timestamps, and reference audio, seamlessly combining speech, sound effects, ambient sounds, and more. The model supports a wide range of tasks, including multilingual TTS, single-speaker speech, multi-speaker dialogue, podcasts, cinematic soundscapes, ambient audio, and sound effects. Its key strengths include more natural and expressive speech, consistent voice identity, flexible timing control, and high-quality soundscape generation. These capabilities make it well suited for content creation, short-form video, podcasts, film and television production, game audio, and other professional audio creation scenarios.

Inference service provider

Alibaba Cloud Model Studio provides inference services for this model.

Model capabilities

CapabilitySupportCapabilitySupport
Input modalitiesText and reference audioOutput modalityAudio
LanguagesChinese and EnglishOutput modesNon-streaming
Output formatsWAV, MP3, and PCMReference audioURL or Base64

Context limits

ParameterValue
Maximum input length3,000 characters
Maximum output duration per requestPodcasts: 240 seconds (4 minutes); other scenarios: 120 seconds
Reference audio countUp to 3 clips
Duration per reference clipUp to 30 seconds
Size per reference clipUp to 10 MB
Reference audio formatsWAV, MP3, and OGG Opus. Raw PCM is not supported.

Model pricing

The following list price excludes promotions. For promotional offers, see the Model Studio console.
  • China (Beijing)
Billing itemUnit price (USD/million tokens)
Input0.848
Output1.696
Input and output tokens are billed separately. For details, see Model pricing.

Rate limits

  • China (Beijing)
MetricValue
Requests per second (RPS)3

API access

Use the HTTPS API to receive complete audio files. For examples and prompt guidance, see Audio generation. For parameters and sample code, see Audio generation API reference.
Token Plan
Model Playground
  • Audio generation
  • Music generation
Statistics and Monitoring
Asset Center
Support