Qwen3-ASR-Flash is a highly accurate, intelligent, and robust multilingual speech recognition model based on a large language model. Leveraging a powerful foundational model, massive amounts of text and multimodal data, and tens of millions of hours of audio data, Qwen3-ASR-Flash achieves high-precision speech recognition. It can automatically determine the language and accurately recognize speech in multiple languages, ensuring precise transcription even in complex audio environments.This model version is functionally equivalent to the snapshot model qwen3-asr-flash-2025-09-08.
Inference Service Provider
The inference service provider for qwen3-asr-flash is Alibaba Cloud Model Studio.
Model Capabilities
| Capability | Support | Capability | Support |
|---|---|---|---|
Input Modality | Audio | Output Modality | Text |
Model Experience | Unsupported | Function Calling | Unsupported |
Structured Outputs | Unsupported | Web Search | Unsupported |
Prefix Completion | Unsupported | Context Caching | Unsupported |
Batch Inference | Unsupported | Fine-tuning | Unsupported |
Context Limits
| Parameter | Value | Parameter | Value |
|---|---|---|---|
Max Input Length | — | Max Output Length | — |
Context Window | — |
Pricing
This page only shows the original pricing for model API calls, excluding any limited-time promotions. Visit Model Studio Console for promotional offers.
- China (Beijing)
- Singapore
- US (Virginia)
| Billing Item | Price (USD) | Unit |
|---|---|---|
Audio Duration | 0.000032 | Per second |
Rate Limits
- China (Beijing)
- Singapore
- US (Virginia)
| Parameter | Value |
|---|---|
RPM (Requests Per Minute) | 100 |
Snapshot Versions
qwen3-asr-flash-2026-02-10
Qwen3-ASR-Flash is a highly accurate, intelligent, and robust multilingual speech recognition model based on a large language model. Leveraging a powerful foundational model, massive amounts of text and multimodal data, and tens of millions of hours of audio data, Qwen3-ASR-Flash achieves highly accurate speech recognition, automatically determining the language and accurately identifying speech in multiple languages, while ensuring precise transcription even in complex audio environments.This version is a snapshot dated February 10, 2026.
Inference Service Provider
The inference service provider for qwen3-asr-flash-2026-02-10 is Alibaba Cloud Model Studio.
Model Capabilities
| Capability | Support | Capability | Support |
|---|---|---|---|
Input Modality | Audio | Output Modality | Text |
Model Experience | Unsupported | Function Calling | Unsupported |
Structured Outputs | Unsupported | Web Search | Unsupported |
Prefix Completion | Unsupported | Context Caching | Unsupported |
Batch Inference | Unsupported | Fine-tuning | Unsupported |
Context Limits
| Parameter | Value | Parameter | Value |
|---|---|---|---|
Max Input Length | — | Max Output Length | — |
Context Window | — |
Pricing
This page only shows the original pricing for model API calls, excluding any limited-time promotions. Visit Model Studio Console for promotional offers.
- China (Beijing)
- Singapore
| Billing Item | Price (USD) | Unit |
|---|---|---|
Audio Duration | 0.000032 | Per second |
Rate Limits
- China (Beijing)
- Singapore
| Parameter | Value |
|---|---|
RPM (Requests Per Minute) | 100 |
qwen3-asr-flash-2025-09-08
Qwen3-ASR-Flash is a highly accurate, intelligent, and robust multilingual speech recognition model based on a large language model. Leveraging a powerful foundational model, massive amounts of text and multimodal data, and tens of millions of hours of audio data, Qwen3-ASR-Flash achieves high-precision speech recognition. It can automatically determine the language and accurately recognize speech in multiple languages, ensuring precise transcription even in complex audio environments.This version is a snapshot version from September 8, 2025.
Inference Service Provider
The inference service provider for qwen3-asr-flash-2025-09-08 is Alibaba Cloud Model Studio.
Model Capabilities
| Capability | Support | Capability | Support |
|---|---|---|---|
Input Modality | Audio | Output Modality | Text |
Model Experience | Unsupported | Function Calling | Unsupported |
Structured Outputs | Unsupported | Web Search | Unsupported |
Prefix Completion | Unsupported | Context Caching | Unsupported |
Batch Inference | Unsupported | Fine-tuning | Unsupported |
Context Limits
| Parameter | Value | Parameter | Value |
|---|---|---|---|
Max Input Length | — | Max Output Length | — |
Context Window | — |
Pricing
This page only shows the original pricing for model API calls, excluding any limited-time promotions. Visit Model Studio Console for promotional offers.
- China (Beijing)
- Singapore
- US (Virginia)
| Billing Item | Price (USD) | Unit |
|---|---|---|
Audio Duration | 0.000032 | Per second |
Rate Limits
- China (Beijing)
- Singapore
- US (Virginia)
| Parameter | Value |
|---|---|
RPM (Requests Per Minute) | 100 |