The real-time version of Qwen's new large multimodal understanding and generation model, suitable for real-time audio interaction scenarios. It supports the understanding of audio accompanied by text, images, and video mixed inputs, and can simultaneously generate speech and text in stream, providing four natural tones.This model version is functionally equivalent to the snapshot model qwen-omni-turbo-realtime-2025-05-08.
Inference Service Provider
The inference service provider for qwen-omni-turbo-realtime is Alibaba Cloud Model Studio.
Model Capabilities
| Capability | Support | Capability | Support |
|---|---|---|---|
Input Modality | Text Image Video Audio | Output Modality | Text Audio |
Model Experience | Unsupported | Function Calling | Unsupported |
Structured Outputs | Unsupported | Web Search | Unsupported |
Prefix Completion | Unsupported | Context Caching | Unsupported |
Batch Inference | Unsupported | Fine-tuning | Unsupported |
Context Limits
| Parameter | Value | Parameter | Value |
|---|---|---|---|
Max Input Length | 30720 | Max Output Length | 2048 |
Context Window | 32768 |
Pricing
This page only shows the original pricing for model API calls, excluding any limited-time promotions. Visit Model Studio Console for promotional offers.
- China (Beijing)
- Singapore
| Billing Item | Price (USD) | Unit |
|---|---|---|
Input: Text | 0.23 | Per 1M tokens |
Input: Audio | 3.584 | Per 1M tokens |
Input: Vision | 0.861 | Per 1M tokens |
Output: Text (When input contains only text) | 0.918 | Per 1M tokens |
Output: Text (When input contains images/audio/video) | 2.581 | Per 1M tokens |
Output: Text&Audio (Output text is not charged) | 7.168 | Per 1M tokens |
Rate Limits
- China (Beijing)
- Singapore
| Parameter | Value |
|---|---|
RPM (Requests Per Minute) | 60 |
TPM (Tokens Per Minute) | 100,000 |
Dynamic Updates
qwen-omni-turbo-realtime-latest
The real-time version of Qwen's new large multimodal understanding and generation model. This model is a dynamically updated version.
Inference Service Provider
The inference service provider for qwen-omni-turbo-realtime-latest is Alibaba Cloud Model Studio.
Model Capabilities
| Capability | Support | Capability | Support |
|---|---|---|---|
Input Modality | Text Image Video Audio | Output Modality | Text Audio |
Model Experience | Unsupported | Function Calling | Unsupported |
Structured Outputs | Unsupported | Web Search | Unsupported |
Prefix Completion | Unsupported | Context Caching | Unsupported |
Batch Inference | Unsupported | Fine-tuning | Unsupported |
Context Limits
| Parameter | Value | Parameter | Value |
|---|---|---|---|
Max Input Length | 30720 | Max Output Length | 2048 |
Context Window | 32768 |
Pricing
This page only shows the original pricing for model API calls, excluding any limited-time promotions. Visit Model Studio Console for promotional offers.
- China (Beijing)
- Singapore
| Billing Item | Price (USD) | Unit |
|---|---|---|
Input: Text | 0.23 | Per 1M tokens |
Input: Audio | 3.584 | Per 1M tokens |
Input: Vision | 0.861 | Per 1M tokens |
Output: Text (When input contains only text) | 0.918 | Per 1M tokens |
Output: Text (When input contains images/audio/video) | 2.581 | Per 1M tokens |
Output: Text&Audio (Output text is not charged) | 7.168 | Per 1M tokens |
Rate Limits
- China (Beijing)
- Singapore
| Parameter | Value |
|---|---|
RPM (Requests Per Minute) | 60 |
TPM (Tokens Per Minute) | 100,000 |
Snapshot Versions
qwen-omni-turbo-realtime-2025-05-08
The real-time version of Qwen's new large multimodal understanding and generation model, suitable for real-time audio interaction scenarios. It supports the understanding of audio accompanied by text, images, and video mixed inputs, and can simultaneously generate speech and text in stream, providing four natural tones. This model is the snapshot from May 8, 2025.
Inference Service Provider
The inference service provider for qwen-omni-turbo-realtime-2025-05-08 is Alibaba Cloud Model Studio.
Model Capabilities
| Capability | Support | Capability | Support |
|---|---|---|---|
Input Modality | Text Image Video Audio | Output Modality | Text Audio |
Model Experience | Unsupported | Function Calling | Unsupported |
Structured Outputs | Unsupported | Web Search | Unsupported |
Prefix Completion | Unsupported | Context Caching | Unsupported |
Batch Inference | Unsupported | Fine-tuning | Unsupported |
Context Limits
| Parameter | Value | Parameter | Value |
|---|---|---|---|
Max Input Length | 30720 | Max Output Length | 2048 |
Context Window | 32768 |
Pricing
This page only shows the original pricing for model API calls, excluding any limited-time promotions. Visit Model Studio Console for promotional offers.
- China (Beijing)
- Singapore
| Billing Item | Price (USD) | Unit |
|---|---|---|
Input: Text | 0.23 | Per 1M tokens |
Input: Audio | 3.584 | Per 1M tokens |
Input: Vision | 0.861 | Per 1M tokens |
Output: Text (When input contains only text) | 0.918 | Per 1M tokens |
Output: Text (When input contains images/audio/video) | 2.581 | Per 1M tokens |
Output: Text&Audio (Output text is not charged) | 7.168 | Per 1M tokens |
Rate Limits
- China (Beijing)
- Singapore
| Parameter | Value |
|---|---|
RPM (Requests Per Minute) | 60 |
TPM (Tokens Per Minute) | 100,000 |