Skip to main content
Speech Recognition

qwen-audio-3.1-asr-flash-message

Qwen-Audio-3.1-ASR-Flash-Message is a generative speech recognition model optimized for voice messages, voice input, and enterprise scenarios. It supports Chinese, English, and other languages, with optimizations for Indonesian, Japanese, Thai, Filipino, Vietnamese, Korean, and Malay. The model supports incremental streaming transcription, P0/P1 hotwords, text normalization, recognition informed by preceding transcripts and business context, and optional native text polishing. It improves recognition stability with multiple speakers, background conversations, far-field audio, and environmental noise. It supports automatic language detection or specified target languages, and can transcribe Chinese dialects as spoken or convert them to standard Mandarin text.

Inference Service Provider

The inference service provider for qwen-audio-3.1-asr-flash-message is Alibaba Cloud Model Studio.

Model Capabilities

CapabilitySupportCapabilitySupport
Input ModalityAudioOutput ModalityText
Model ExperienceUnsupportedFunction CallingUnsupported
Structured OutputsSupportedWeb SearchUnsupported
Prefix CompletionUnsupportedContext CachingUnsupported
Batch InferenceUnsupportedFine-tuningUnsupported

Context Limits

ParameterValueParameterValue
Max Input Length7168 TokenMax Output Length1024 Token
Context Window8192 Token

Pricing

This page only shows the original pricing for model API calls, excluding any limited-time promotions. Visit Model Studio Console for promotional offers.
  • China (Beijing)
  • Singapore
Billing itemPrice (USD)Unit
Input0.848Per million tokens
Output0.636Per million tokens

Rate Limits

  • China (Beijing)
  • Singapore
ParameterValue
RPM (requests per minute)1200

API usage

Token Plan
Model Playground
  • Audio generation
  • Music generation
Statistics and Monitoring