Skip to main content
Omni-modal

qwen2.5-omni-7b

The new multi-modal understanding and generation model trained based on Qwen2.5. It supports text, image, speech, video, and mixed input understanding and can simultaneously generate streams of text and speech, significantly improves the speed of multi-modal content understanding. It provides four natural tones.

Inference Service Provider

The inference service provider for qwen2.5-omni-7b is Alibaba Cloud Model Studio.

Model Capabilities

CapabilitySupportCapabilitySupport

Input Modality

Text Image Video Audio

Output Modality

Text Audio

Model Experience

Supported

Function Calling

Unsupported

Structured Outputs

Unsupported

Web Search

Unsupported

Prefix Completion

Unsupported

Context Caching

Unsupported

Batch Inference

Unsupported

Fine-tuning

Unsupported

Context Limits

ParameterValueParameterValue

Max Input Length

30720

Max Output Length

2048

Context Window

32768

Pricing

This page only shows the original pricing for model API calls, excluding any limited-time promotions. Visit Model Studio Console for promotional offers.
  • China (Beijing)
  • Singapore
Billing ItemPrice (USD)Unit

Input: Text

0.087

Per 1M tokens

Input: Audio

5.448

Per 1M tokens

Input: Vision

0.287

Per 1M tokens

Output: Text (When input contains only text)

0.345

Per 1M tokens

Output: Text (When input contains images/audio/video)

0.861

Per 1M tokens

Output: Text&Audio (Output text is not charged)

10.895

Per 1M tokens

Rate Limits

  • China (Beijing)
  • Singapore
ParameterValue

RPM (Requests Per Minute)

60

TPM (Tokens Per Minute)

100,000

Token Plan
Model Playground
  • Audio generation
  • Music generation
Statistics and Monitoring
Asset Center
Support