Skip to main content
Vision

qwen-vl-plus

Qwen-VL-Plus is the enhanced version of the large visual language model. It significantly improves detail recognition and text recognition capabilities, supporting images with resolutions exceeding one million pixels and any aspect ratio specifications. The model delivers exceptional performance across a wide range of visual tasks.This model version is functionally equivalent to the snapshot model qwen-vl-plus-2025-08-15.

Inference Service Provider

The inference service provider for qwen-vl-plus is Alibaba Cloud Model Studio.

Model Capabilities

  • China (Beijing)
  • Singapore
CapabilitySupportCapabilitySupport

Input Modality

Text Image Video

Output Modality

Text

Model Experience

Supported

Function Calling

Unsupported

Structured Outputs

Supported

Web Search

Unsupported

Prefix Completion

Supported

Context Caching

Supported

Batch Inference

Supported

Fine-tuning

Supported

Context Limits

ParameterValueParameterValue

Max Input Length

129024

Max Output Length

8192

Context Window

131072

Pricing

This page only shows the original pricing for model API calls, excluding any limited-time promotions. Visit Model Studio Console for promotional offers.
  • China (Beijing)
  • Singapore
Billing ItemPrice (USD)Unit

Input

0.115

Per 1M tokens

Output

0.287

Per 1M tokens

Input(Implicit Cache)

0.023

Per 1M tokens

Input(Batch File)

0.057

Per 1M tokens

Output(Batch File)

0.143

Per 1M tokens

Rate Limits

  • China (Beijing)
  • Singapore
ParameterValue

RPM (Requests Per Minute)

1200

TPM (Tokens Per Minute)

1,000,000

Token Plan
Model Playground
  • Music generation
Statistics and Monitoring