Skip to main content
Omni-modal

qwen3-omni-flash-realtime

The real-time version of the Qwen3-Omni-Flash multimodal large-scale model, based on the Thinker–Talker Mixed Expert (MoE) architecture, supports efficient understanding and speech generation of text, images, audio, and video. It can interact with text in 119 languages ​​and speech in 20 languages, generating human-like speech for precise cross-lingual communication. The model boasts powerful command-following and system prompt customization capabilities, flexibly adapting to conversational styles and character settings. It is widely used in scenarios such as text creation, voice assistants, and multimedia analysis, providing a natural and smooth multimodal interaction experience.This model version is functionally equivalent to the snapshot model qwen3-omni-flash-realtime-2025-12-01.

Inference Service Provider

The inference service provider for qwen3-omni-flash-realtime is Alibaba Cloud Model Studio.

Model Capabilities

CapabilitySupportCapabilitySupport

Input Modality

Text Image Video Audio

Output Modality

Text Audio

Model Experience

Unsupported

Function Calling

Unsupported

Structured Outputs

Unsupported

Web Search

Unsupported

Prefix Completion

Unsupported

Context Caching

Unsupported

Batch Inference

Unsupported

Fine-tuning

Unsupported

Context Limits

ParameterValueParameterValue

Max Input Length

49152

Max Output Length

16384

Context Window

65536

Pricing

This page only shows the original pricing for model API calls, excluding any limited-time promotions. Visit Model Studio Console for promotional offers.
  • China (Beijing)
  • Singapore
Billing ItemPrice (USD)Unit

Input: Text

0.315

Per 1M tokens

Input: Audio

2.709

Per 1M tokens

Input: Vision

0.559

Per 1M tokens

Output: Text (When input contains only text)

1.19

Per 1M tokens

Output: Text (When input contains images/audio/video)

2.179

Per 1M tokens

Output: Text&Audio (Output text is not charged)

10.766

Per 1M tokens

Rate Limits

  • China (Beijing)
  • Singapore
ParameterValue

RPM (Requests Per Minute)

60

TPM (Tokens Per Minute)

100,000

Snapshot Versions

qwen3-omni-flash-realtime-2025-12-01

This is the real-time version of Qwen3-Omni-Flash multimodal large model, based on the Thinker-Talker Hybrid Expert (MoE) architecture. It supports efficient understanding and speech generation of text, images, audio, and video, enabling text interaction in 119 languages ​​and voice interaction in 20 languages. It supports 49 voice timbres and generates human-like speech for accurate cross-language communication. The model features powerful command following and system prompt customization capabilities, flexibly adapting to dialogue styles and role settings. It is widely used in text creation, voice assistants, multimedia analysis, and other scenarios, providing a natural and smooth multimodal interactive experience. This version is a snapshot from December 1, 2025.

Inference Service Provider

The inference service provider for qwen3-omni-flash-realtime-2025-12-01 is Alibaba Cloud Model Studio.

Model Capabilities

CapabilitySupportCapabilitySupport

Input Modality

Text Image Audio Video

Output Modality

Text Audio

Model Experience

Unsupported

Function Calling

Unsupported

Structured Outputs

Unsupported

Web Search

Unsupported

Prefix Completion

Unsupported

Context Caching

Unsupported

Batch Inference

Unsupported

Fine-tuning

Unsupported

Context Limits

ParameterValueParameterValue

Max Input Length

49152

Max Output Length

16384

Context Window

65536

Pricing

This page only shows the original pricing for model API calls, excluding any limited-time promotions. Visit Model Studio Console for promotional offers.
  • China (Beijing)
Billing ItemPrice (USD)Unit

Input: Text

0.315

Per 1M tokens

Input: Audio

2.709

Per 1M tokens

Input: Vision

0.559

Per 1M tokens

Output: Text (When input contains only text)

1.19

Per 1M tokens

Output: Text (When input contains images/audio/video)

2.179

Per 1M tokens

Output: Text&Audio (Output text is not charged)

10.766

Per 1M tokens

Rate Limits

  • China (Beijing)
  • Singapore
ParameterValue

RPM (Requests Per Minute)

60

TPM (Tokens Per Minute)

100,000

qwen3-omni-flash-realtime-2025-09-15

The real-time version of the Qwen3-Omni-Flash multimodal large-scale model, based on the Thinker–Talker Mixed Expert (MoE) architecture, supports efficient understanding and speech generation of text, images, audio, and video. It can interact with text in 119 languages ​​and speech in 20 languages, generating human-like speech for precise cross-lingual communication. The model boasts powerful command-following and system prompt customization capabilities, flexibly adapting to conversational styles and character settings. It is widely used in scenarios such as text creation, voice assistants, and multimedia analysis, providing a natural and smooth multimodal interaction experience.

Inference Service Provider

The inference service provider for qwen3-omni-flash-realtime-2025-09-15 is Alibaba Cloud Model Studio.

Model Capabilities

  • China (Beijing)
  • Singapore
CapabilitySupportCapabilitySupport

Input Modality

Text Image Video Audio

Output Modality

Text Audio

Model Experience

Unsupported

Function Calling

Unsupported

Structured Outputs

Unsupported

Web Search

Unsupported

Prefix Completion

Unsupported

Context Caching

Unsupported

Batch Inference

Unsupported

Fine-tuning

Unsupported

Context Limits

ParameterValueParameterValue

Max Input Length

49152

Max Output Length

16384

Context Window

65536

Pricing

This page only shows the original pricing for model API calls, excluding any limited-time promotions. Visit Model Studio Console for promotional offers.
  • China (Beijing)
  • Singapore
Billing ItemPrice (USD)Unit

Input: Text

0.315

Per 1M tokens

Input: Audio

2.709

Per 1M tokens

Input: Vision

0.559

Per 1M tokens

Output: Text (When input contains only text)

1.19

Per 1M tokens

Output: Text (When input contains images/audio/video)

2.179

Per 1M tokens

Output: Text&Audio (Output text is not charged)

10.766

Per 1M tokens

Rate Limits

  • China (Beijing)
  • Singapore
ParameterValue

RPM (Requests Per Minute)

60

TPM (Tokens Per Minute)

100,000

Token Plan
Model Playground
  • Audio generation
  • Music generation
Statistics and Monitoring
Asset Center
Support