Skip to main content
Omni-modal

qwen-omni-turbo

The new multimodal understanding and generation model, supporting text, image, speech, and video input as well as mixed input understanding. It can simultaneously stream text and speech, with significantly enhanced speed for multimodal content understanding. Additionally, it offers four natural dialogue tones.This model version is functionally equivalent to the snapshot model qwen-omni-turbo-2025-03-26.

Inference Service Provider

The inference service provider for qwen-omni-turbo is Alibaba Cloud Model Studio.

Model Capabilities

  • China (Beijing)
  • Singapore
CapabilitySupportCapabilitySupport

Input Modality

Text Image Video Audio

Output Modality

Text Audio

Model Experience

Supported

Function Calling

Unsupported

Structured Outputs

Unsupported

Web Search

Unsupported

Prefix Completion

Unsupported

Context Caching

Unsupported

Batch Inference

Supported

Fine-tuning

Unsupported

Context Limits

ParameterValueParameterValue

Max Input Length

30720

Max Output Length

2048

Context Window

32768

Pricing

This page only shows the original pricing for model API calls, excluding any limited-time promotions. Visit Model Studio Console for promotional offers.
  • China (Beijing)
  • Singapore
Billing ItemPrice (USD)Unit

Input: Text

0.058

Per 1M tokens

Input: Audio

3.584

Per 1M tokens

Input: Vision

0.216

Per 1M tokens

Output: Text (When input contains only text)

0.23

Per 1M tokens

Output: Text (When input contains images/audio/video)

0.646

Per 1M tokens

Output: Text&Audio (Output text is not charged)

7.168

Per 1M tokens

Input: Vision(Implicit Cache)

0.044

Per 1M tokens

Input: Text(Implicit Cache)

0.012

Per 1M tokens

Input: Audio(Implicit Cache)

0.717

Per 1M tokens

Rate Limits

  • China (Beijing)
  • Singapore
ParameterValue

RPM (Requests Per Minute)

60

TPM (Tokens Per Minute)

100,000

Dynamic Updates

qwen-omni-turbo-latest

The new multimodal understanding and generation model, supporting text, image, speech, and video input as well as mixed input understanding. It can simultaneously stream text and speech, with significantly enhanced speed for multimodal content understanding. Additionally, it offers four natural dialogue tones. This model is a dynamically updated version.

Inference Service Provider

The inference service provider for qwen-omni-turbo-latest is Alibaba Cloud Model Studio.

Model Capabilities

CapabilitySupportCapabilitySupport

Input Modality

Text Image Video Audio

Output Modality

Text Audio

Model Experience

Supported

Function Calling

Unsupported

Structured Outputs

Unsupported

Web Search

Unsupported

Prefix Completion

Unsupported

Context Caching

Unsupported

Batch Inference

Unsupported

Fine-tuning

Unsupported

Context Limits

ParameterValueParameterValue

Max Input Length

30720

Max Output Length

2048

Context Window

32768

Pricing

This page only shows the original pricing for model API calls, excluding any limited-time promotions. Visit Model Studio Console for promotional offers.
  • China (Beijing)
  • Singapore
Billing ItemPrice (USD)Unit

Input: Text

0.058

Per 1M tokens

Input: Audio

3.584

Per 1M tokens

Input: Vision

0.216

Per 1M tokens

Output: Text (When input contains only text)

0.23

Per 1M tokens

Output: Text (When input contains images/audio/video)

0.646

Per 1M tokens

Output: Text&Audio (Output text is not charged)

7.168

Per 1M tokens

Rate Limits

  • China (Beijing)
  • Singapore
ParameterValue

RPM (Requests Per Minute)

60

TPM (Tokens Per Minute)

100,000

Snapshot Versions

qwen-omni-turbo-2025-03-26

The new multimodal understanding and generation model, supporting text, image, speech, and video input as well as mixed input understanding. It can simultaneously stream text and speech, with significantly enhanced speed for multimodal content understanding. Additionally, it offers four natural dialogue tones. This model is the snapshot from March 26, 2025, with significant improvements on visual capabilities over the snapshot of January 19, 2025.

Inference Service Provider

The inference service provider for qwen-omni-turbo-2025-03-26 is Alibaba Cloud Model Studio.

Model Capabilities

CapabilitySupportCapabilitySupport

Input Modality

Text Image Video Audio

Output Modality

Text Audio

Model Experience

Supported

Function Calling

Unsupported

Structured Outputs

Unsupported

Web Search

Unsupported

Prefix Completion

Unsupported

Context Caching

Unsupported

Batch Inference

Unsupported

Fine-tuning

Unsupported

Context Limits

ParameterValueParameterValue

Max Input Length

30720

Max Output Length

2048

Context Window

32768

Pricing

This page only shows the original pricing for model API calls, excluding any limited-time promotions. Visit Model Studio Console for promotional offers.
  • China (Beijing)
  • Singapore
Billing ItemPrice (USD)Unit

Input: Text

0.058

Per 1M tokens

Input: Audio

3.584

Per 1M tokens

Input: Vision

0.216

Per 1M tokens

Output: Text (When input contains only text)

0.23

Per 1M tokens

Output: Text (When input contains images/audio/video)

0.646

Per 1M tokens

Output: Text&Audio (Output text is not charged)

7.168

Per 1M tokens

Rate Limits

  • China (Beijing)
  • Singapore
ParameterValue

RPM (Requests Per Minute)

60

TPM (Tokens Per Minute)

100,000

qwen-omni-turbo-2025-01-19

Qwen All-Modal Understanding and Generation Large Model supports text, image, speech, video input understanding and mixed input comprehension, features simultaneous streaming generation of text and speech, significantly enhanced multi-modal content understanding speed, provides 4 natural conversational voices. This version is a snapshot from January 19, 2025, and will be maintained until approximately one month before the next snapshot release.

Inference Service Provider

The inference service provider for qwen-omni-turbo-2025-01-19 is Alibaba Cloud Model Studio.

Model Capabilities

CapabilitySupportCapabilitySupport

Input Modality

Text Image Video Audio

Output Modality

Text Audio

Model Experience

Supported

Function Calling

Unsupported

Structured Outputs

Unsupported

Web Search

Unsupported

Prefix Completion

Unsupported

Context Caching

Unsupported

Batch Inference

Unsupported

Fine-tuning

Unsupported

Context Limits

ParameterValueParameterValue

Max Input Length

30720

Max Output Length

2048

Context Window

32768

Pricing

This page only shows the original pricing for model API calls, excluding any limited-time promotions. Visit Model Studio Console for promotional offers.
  • China (Beijing)
Billing ItemPrice (USD)Unit

Input: Text

0.058

Per 1M tokens

Input: Audio

3.584

Per 1M tokens

Input: Vision

0.216

Per 1M tokens

Output: Text (When input contains only text)

0.23

Per 1M tokens

Output: Text (When input contains images/audio/video)

0.646

Per 1M tokens

Output: Text&Audio (Output text is not charged)

7.168

Per 1M tokens

Rate Limits

  • China (Beijing)
ParameterValue

RPM (Requests Per Minute)

60

TPM (Tokens Per Minute)

100,000

Token Plan
Model Playground
  • Audio generation
  • Music generation
Statistics and Monitoring
Asset Center
Support