The new multi-modal understanding and generation model trained based on Qwen2.5. It supports text, image, speech, video, and mixed input understanding and can simultaneously generate streams of text and speech, significantly improves the speed of multi-modal content understanding. It provides four natural tones.
Inference Service Provider
The inference service provider for qwen2.5-omni-7b is Alibaba Cloud Model Studio.
Model Capabilities
| Capability | Support | Capability | Support |
|---|---|---|---|
Input Modality | Text Image Video Audio | Output Modality | Text Audio |
Model Experience | Supported | Function Calling | Unsupported |
Structured Outputs | Unsupported | Web Search | Unsupported |
Prefix Completion | Unsupported | Context Caching | Unsupported |
Batch Inference | Unsupported | Fine-tuning | Unsupported |
Context Limits
| Parameter | Value | Parameter | Value |
|---|---|---|---|
Max Input Length | 30720 | Max Output Length | 2048 |
Context Window | 32768 |
Pricing
This page only shows the original pricing for model API calls, excluding any limited-time promotions. Visit Model Studio Console for promotional offers.
- China (Beijing)
- Singapore
| Billing Item | Price (USD) | Unit |
|---|---|---|
Input: Text | 0.087 | Per 1M tokens |
Input: Audio | 5.448 | Per 1M tokens |
Input: Vision | 0.287 | Per 1M tokens |
Output: Text (When input contains only text) | 0.345 | Per 1M tokens |
Output: Text (When input contains images/audio/video) | 0.861 | Per 1M tokens |
Output: Text&Audio (Output text is not charged) | 10.895 | Per 1M tokens |
Rate Limits
- China (Beijing)
- Singapore
| Parameter | Value |
|---|---|
RPM (Requests Per Minute) | 60 |
TPM (Tokens Per Minute) | 100,000 |