Qwen 3.5-Omni is a Qwen multimodal large model, supporting text, image, audio, and audio-visual understanding and interaction. As a fully evolved version of Qwen3-Omni, it supports audio input in 60+ languages, voice output in 30+ languages, and controllable voice dialogue, WebSearch and complex FunctionCall invocation, and has intelligent semantic interruption interaction capabilities. It is widely used in scenarios such as text creation, voice assistants, and multimedia analysis, providing a natural and smooth multimodal interactive experience.This model version is functionally equivalent to the snapshot model qwen3.5-omni-plus-realtime-2026-03-15.
Inference Service Provider
The inference service provider for qwen3.5-omni-plus-realtime is Alibaba Cloud Model Studio.
Model Capabilities
| Capability | Support | Capability | Support |
|---|---|---|---|
Input Modality | Text Image Video Audio | Output Modality | Text Audio |
Model Experience | Unsupported | Function Calling | Supported |
Structured Outputs | Unsupported | Web Search | Supported |
Prefix Completion | Unsupported | Context Caching | Unsupported |
Batch Inference | Unsupported | Fine-tuning | Unsupported |
Context Limits
| Parameter | Value | Parameter | Value |
|---|---|---|---|
Max Input Length | 196608 | Max Output Length | 65536 |
Context Window | 262144 |
Pricing
This page only shows the original pricing for model API calls, excluding any limited-time promotions. Visit Model Studio Console for promotional offers.
- China (Beijing)
- Singapore
| Billing Item | Price (USD) | Unit |
|---|---|---|
Input: Audio | 11 | Per 1M tokens |
Output: Text&Audio (Output text is not charged) | 41.26 | Per 1M tokens |
input:Text/Image/Video | 1.38 | Per 1M tokens |
Output: Text | 8.25 | Per 1M tokens |
Rate Limits
- China (Beijing)
- Singapore
| Parameter | Value |
|---|---|
RPM (Requests Per Minute) | 60 |
TPM (Tokens Per Minute) | 100,000 |
Snapshot Versions
qwen3.5-omni-plus-realtime-2026-03-15
Qwen 3.5-Omni is a Qwen multimodal large model, supporting text, image, audio, and audio-visual understanding and interaction. As a fully evolved version of Qwen3-Omni, it supports audio input in 60+ languages, voice output in 30+ languages, and controllable voice dialogue, WebSearch and complex FunctionCall invocation, and has intelligent semantic interruption interaction capabilities. It is widely used in scenarios such as text creation, voice assistants, and multimedia analysis, providing a natural and smooth multimodal interactive experience.This version is a snapshot from March 15, 2026.
Inference Service Provider
The inference service provider for qwen3.5-omni-plus-realtime-2026-03-15 is Alibaba Cloud Model Studio.
Model Capabilities
| Capability | Support | Capability | Support |
|---|---|---|---|
Input Modality | Text Image Video Audio | Output Modality | Text Audio |
Model Experience | Unsupported | Function Calling | Supported |
Structured Outputs | Unsupported | Web Search | Supported |
Prefix Completion | Unsupported | Context Caching | Unsupported |
Batch Inference | Unsupported | Fine-tuning | Unsupported |
Context Limits
| Parameter | Value | Parameter | Value |
|---|---|---|---|
Max Input Length | 196608 | Max Output Length | 65536 |
Context Window | 262144 |
Pricing
This page only shows the original pricing for model API calls, excluding any limited-time promotions. Visit Model Studio Console for promotional offers.
- China (Beijing)
- Singapore
| Billing Item | Price (USD) | Unit |
|---|---|---|
Input: Audio | 11 | Per 1M tokens |
Output: Text&Audio (Output text is not charged) | 41.26 | Per 1M tokens |
input:Text/Image/Video | 1.38 | Per 1M tokens |
Output: Text | 8.25 | Per 1M tokens |
Rate Limits
- China (Beijing)
- Singapore
| Parameter | Value |
|---|---|
RPM (Requests Per Minute) | 60 |
TPM (Tokens Per Minute) | 100,000 |