The Qwen-Omni model accepts multimodal input and generates text or speech responses. It produces human-like voices and supports speech output in multiple languages and dialects. Use cases include content moderation, text creation, visual recognition, and audio-video interaction.
Supported regions:Singapore, Beijing. Use the API key for your region.
Prerequisites
All requests to Qwen-Omni must set
Configure parameters, prompts, and media lengths to balance cost, speed, and quality.
The following example shows how to provide an image and audio in a single request for multimodal analysis.
Each request contains text and one other modality (video, audio, or image). All Qwen-Omni models support this.
The Qwen3.5-Omni series supports web search to retrieve real-time information and perform reasoning.
In the Qwen-Omni series, only the Qwen3-Omni-Flash model is a hybrid thinking model. You can use the
When using Qwen-Omni models for multi-turn conversations, note the following:
Qwen-Omni models output audio as streaming Base64-encoded data. During generation, maintain a string variable and append the Base64-encoded data from each returned chunk. After generation completes, Base64-decode the complete string to get the audio file. Alternatively, decode and play each chunk in real time.
For input and output parameters, see OpenAI compatible - Chat.
Billing rules
Qwen-Omni is billed based on tokens consumed across modalities (audio, image, and video). Check billing details in the console.
Free quota
To claim, query, or use your free quota, see Free quota for new users.
Rate limits
For rate limiting rules and FAQ, see Rate limiting.
If the model call fails and returns an error message, see Error codes for resolution.
For a list of voices for the Qwen-Omni model, see Voice list.
Getting started
Prerequisites
- Obtain an API key and set the API key as an environment variable.
- The Qwen-Omni model supports only OpenAI-compatible invocation. You must install the latest SDK. The minimum required versions are 1.52.0 for the OpenAI Python SDK and 4.68.0 for the Node.js SDK.
Response
Response
After you run the Running
Python or Node.js code, the text response appears in the console and an audio file named audio_assistant.wav is saved in the same directory as your code file.HTTP code returns text and Base64-encoded audio data directly in the audio field.Model selection
-
Qwen3.5-Omni series: Best for long video analysis, meeting summaries, caption generation, content moderation, and audio-video interaction.
- Input limits: Up to 3 hours of audio or 1 hour of video
- Audio control: Supports adjusting volume, speaking rate, and emotion through instructions
- Visual capability: Matches the level of Qwen3.5. Understands images, speech, sound effects, and other multimodal input
- Combined multimodal input: Supports any combination of text with images, audio, and video in a single request
- Voice cloning: Supports custom voices (only qwen3.5-omni-plus and qwen3.5-omni-flash; snapshot versions are not supported). For details, see Voice cloning
-
Qwen3-Omni-Flash series: Best for short video analysis and cost-sensitive scenarios.
- Input limits: Audio and video input up to 150 seconds
- Thinking mode: The only Qwen-Omni series model that supports thinking mode
- Input modality: Supports only a combination of text with a single other modality (image, audio, or video).
- Qwen-Omni-Turbo series This series is no longer updated and has limited features. We recommend migrating to the Qwen3.5-Omni or Qwen3-Omni-Flash series.
| Series | Audio-video description | Deep thinking | Web search | Input audio languages | Output audio languages | Supported voices |
| Qwen3.5-OmniLatest-generation omni-modal model | Strong | Not supported | Supported | 113
74 languages and 39 dialects Languages: Chinese, English, German, French, Italian, Czech, Indonesian, Thai, Korean, Polish, Japanese, Vietnamese, Finnish, Portuguese, Spanish, Dutch, Russian, Malay, Catalan, Swedish, Turkish, Ukrainian, Romanian, Slovak, Danish, Icelandic, Norwegian (Bokmål), Macedonian, Greek, Hungarian, Galician, Filipino, Croatian, Bosnian, Slovenian, Bulgarian, Kazakh, Belarusian, Latvian, Estonian, Azerbaijani, Uyghur, Swahili, Hindi, Esperanto, Kyrgyz, Tajik, Cebuano, Afrikaans, Arabic, Lithuanian, Javanese, Bengali, Persian, Hebrew, Punjabi, Gujarati, Mongolian, Asturian, Kannada, Marathi, Interlingua, Malayalam, Maltese, Norwegian Nynorsk, Telugu, Urdu, Georgian, Basque, Tamil, Odia, Serbian, MaoriDialects: Northeastern Mandarin, Guizhou dialect, Cantonese, Henan dialect, Hong Kong Cantonese, Shanghainese, Shaanxi dialect, Tianjin dialect, Taiwanese Hokkien, Yunnan dialect, Anhui dialect, Fujian dialect, Gansu dialect, Guangdong dialect, Hubei dialect, Hunan dialect, Jiangxi dialect, Shandong dialect, Shanxi dialect, Sichuan dialect, Guangxi dialect, Hainan dialect, Chongqing dialect, Changsha dialect, Hangzhou dialect, Hefei dialect, Yinchuan dialect, Zhengzhou dialect, Shenyang dialect, Wenzhou dialect, Wuhan dialect, Kunming dialect, Taiyuan dialect, Nanchang dialect, Jinan dialect, Lanzhou dialect, Nanjing dialect, Hakka, Southern Min | 36
29 languages and 7 dialects Languages:Chinese, English, German, Italian, Portuguese, Spanish, Japanese, Korean, French, Russian, Thai, Indonesian, Arabic, Vietnamese, Turkish, Finnish, Polish, Hindi, Dutch, Czech, Urdu, Tagalog, Swedish, Danish, Hebrew, Icelandic, Malay, Norwegian, Persian Dialects:Sichuan dialect, Beijing dialect, Tianjin dialect, Nanjing dialect, Shaanxi dialect, Cantonese, Southern Min | 55 |
| Qwen3-Omni-FlashHybrid thinking model | Weaker | Supported | Not supported | 19
11 languages and 8 dialects Language:Chinese, English, German, French, Italian, Thai, Korean, Japanese, Russian, Spanish, PortugueseDialects:Sichuan dialect, Shanghainese, Cantonese, Southern Min, Shaanxi dialect, Nanjing dialect, Tianjin dialect, Beijing dialect | 19
11 languages and 8 dialects Language:Chinese, English, German, French, Italian, Thai, Korean, Japanese, Russian, Spanish, PortugueseDialects:Sichuanese, Shanghainese, Cantonese, Hokkien, Shaanxi dialect, Nanjing dialect, Tianjin dialect, Beijing dialect | 17 to 49Varies by version |
| Qwen-Omni-TurboNo longer updated | None | Not supported | Not supported | Chinese, English | Chinese, English | 4 |
For model names, context windows, pricing, and snapshot versions, check in the Model Studio console. For rate limits, see Rate limiting.
Model performance
Audio and video content analysis
Generate a comprehensive timestamped description of this video. | 00:00.000 – 00:02.500A rain-soaked city street fills the widescreen frame. Long-exposure photography creates streaks of red and blue car lights on wet pavement. A solitary man in a dark knee-length trench coat and light shirt with a tie walks toward the camera along the right sidewalk. Raindrops cling to his shoulders and hair. Each step makes a muffled plop on damp concrete. A low, ominous electronic hum underscores ambient sounds, while steady rain crackles all around. Colorful graffiti covers the brick wall beside him, and neon signs glow in the distance—one clearly displays orange-red cursive script reading "Nice," while another vertical sign shows the letters "AT" at its bottom.......00:12.300 – 00:14.533The camera tilts down. Inside a flashlight beam, a yellowed, aged single sheet lies in a shallow puddle, its edges charred and corners curled. Inked paragraphs cover most of its surface, but moisture and scorch marks make the text unreadable. The man kneels, extending gloved fingers toward the document. Ripples spread outward as his hand nears. Sirens peak then begin to fade slightly, while the ambient hum remains steady. The scene freezes just before the trembling paper is touched, ending the clip.Only partial timestamps are shown. |
Multidialect understanding
Describe the vocal characteristics, emotional state, and full lyrics for singers in different Chinese dialects. | This audio features rap performances by singers from various regions of China, showcasing rich dialect diversity. The following are detailed descriptions of each singer's vocal traits, emotional state, and corresponding lyrics:Hangzhou dialect (Wu)Vocal traits:The male singer's voice is soft and gentle, with the distinctive smoothness and warmth of the Wu dialect. His intonation rises and falls gently, and his rhythm is light and quick, creating a relaxed storytelling effect.Emotional state:Relaxed, comfortable, and full of everyday life. He describes Hangzhou's slow-paced lifestyle and urban changes, expressing affection for his hometown and reflections on modern development. Overall, he feels cozy and familiar.Lyrics:> Hangzhou dialect, listen up! Our culture includes local erhua pronunciation.> Like flowers blooming by West Lake, no cultural survey needed—we understand it best.> None of you understand. 'San bu da men', 'ge'r', 'lao'r', 'fen'r', 'ya'r'... how is an outsider supposed to tell any of these apart?> Let's go! Transport is fast—subways everywhere. We'll check it out.> Welcoming Southeast Asia and Thailand, watching "Thirteen Ta" shows. Proud of our city, direct in personality.> So you only came for the "bridge-crossing" experience, and because the flavor is different, you're not returning?Only partial results are shown. |
Lyric caption generation
Transcribe the song lyrics and provide timestamps for each line in this format:[00:00:15,020 --> 00:00:28,085] : When you walk through a storm, hold your head up high.[00:00:28,085 --> 00:00:40,200] And don't be afraid of the dark. ...... | [00:00:12,680 --> 00:00:16,960] Cat thread sways past moonlight on trees.[00:00:18,400 --> 00:00:22,800] Radiators hum 1998 chart hits.[00:00:24.160 → 00:00:28.080] Time parts the mist-like heat waves.[00:00:28,920 --> 00:00:33,000] Neon from the screen shines on my nose bridge.......[00:03:16,720 --> 00:03:21,680] We nestle in the softest ring of the tree trunk.[00:03:22,400 --> 00:03:27,000] Breathing turns residual warmth into honey-sugar.[00:03:28,160 --> 00:03:33,200] The sofa sinks into cloud-fluff shape.[00:03:34,000 --> 00:03:38,800] Every pore soaks in sunshine.[00:04:09,000 --> 00:04:10,020] (End)Only partial results are shown. |
Audio-video programming
Usage
Streaming output
All requests to Qwen-Omni must set stream=True.
Model configuration
Configure parameters, prompts, and media lengths to balance cost, speed, and quality.
- Audio-video understanding
- Audio understanding
| Use case | Recommended video length | Recommended prompt | Recommended max_pixels |
| Fast review, low cost | ≤60 minutes | Simple prompt within 50 words | 230,400 |
| Content extraction (long video segmentation) | ≤60 minutes | 921,600~2,073,600 | |
| Standard analysis (short video tagging) | ≤4 minutes | Use the structured prompt below
Recommended prompt | 921,600~2,073,600 |
| Fine-grained analysis (multiple speakers/complex scenes) | ≤2 minutes | 2,073,600 |
You can segment long videos first to obtain fine-grained descriptions.
Combined multimodal input
Combined multimodal input is supported only by the Qwen3.5-Omni series. You can provide data in multiple modalities, such as any combination of image, audio, and text, or video, image, and text, in the same request.
- OpenAI compatible
Single modality input
Each request contains text and one other modality (video, audio, or image). All Qwen-Omni models support this.
- Video and text input
- Audio and text input
- Image and text input
Provide the video as an image list or a video file (with audio support).
- Video file (supports audio in the video)
- Image list format
-
Number of files:
- Qwen3.5-Omni series: Up to 512 files using public URLs and up to 250 files using Base64 encoding.
- Qwen3-Omni-Flash and Qwen-Omni-Turbo series: Only one file is allowed.
-
File size:
-
Using public URLs:
- Qwen3.5-Omni series: Up to 2 GB
- Qwen3-Omni-Flash: Up to 256 MB
- Qwen-Omni-Turbo: Up to 150 MB
- Using Base64 encoding: The encoded Base64 string must be smaller than 10 MB
-
Using public URLs:
-
Duration limits:
- Qwen3.5-Omni series: 1 hour
- Qwen3-Omni-Flash: 150 seconds
- Qwen-Omni-Turbo: 40 seconds
- File formats: MP4, AVI, MKV, MOV, FLV, and WMV.
- Visual and audio information in the video file are billed separately.
- OpenAI compatible
Web search
The Qwen3.5-Omni series supports web search to retrieve real-time information and perform reasoning.
- Web search is supported only in the Qwen3.5-Omni series. The
search_strategyparameter only acceptsagent. - For billing, see the
agentpolicy in Billing.
enable_search and search_strategy to agent:
- OpenAI compatible
Enable/disable thinking mode
In the Qwen-Omni series, only the Qwen3-Omni-Flash model is a hybrid thinking model. You can use the enable_thinking parameter to enable or disable the thinking mode:
truefalse(default)
In thinking mode, audio output is not supported.
- OpenAI compatible
Response
Response
Multi-turn conversation
When using Qwen-Omni models for multi-turn conversations, note the following:
- Assistant Message Assistant messages in the messages array can contain only text data.
- User Message A user message can contain text and one other modality. In multi-turn conversations, you can input different modalities in different user messages.
- OpenAI compatible
Parsing output Base64-encoded audio data
Qwen-Omni models output audio as streaming Base64-encoded data. During generation, maintain a string variable and append the Base64-encoded data from each returned chunk. After generation completes, Base64-decode the complete string to get the audio file. Alternatively, decode and play each chunk in real time.
Input Base64-encoded local file
When using Base64 encoding to send files, the encoded Base64 string must be smaller than 10 MB.
- Images
- Audio
- Video
This example uses the locally saved file eagle.png.
API reference
For input and output parameters, see OpenAI compatible - Chat.
Billing and rate limits
Billing rules
Qwen-Omni is billed based on tokens consumed across modalities (audio, image, and video). Check billing details in the console.
Token conversion rules for audio, images, and videos
Token conversion rules for audio, images, and videos
- Audio
- Images
- Video
-
Qwen3.5-Omni series:- Input audio formula:
Total tokens = Audio duration (seconds) * 7 - Output audio formula:
Total tokens = Audio duration (seconds) * 12.5
- Input audio formula:
-
Qwen3-Omni-Flash: For both input and output audio,Total tokens = Audio duration (seconds) * 12.5 -
Qwen-Omni-Turbo: For both input and output audio,Total tokens = Audio duration (seconds) * 25