Skip to main content

Visual understanding

Choose the right model for your use case, such as image analysis, video understanding, or OCR.

Image and video understanding

Start with qwen3.7-plus, the flagship Qwen model. It supports 1M context window, up to 2-hour videos, function calling, and built-in tools. Once your application is stable, you can switch to qwen3.7-flash to reduce costs. It offers near-flagship performance with the same context length and feature set.

Image resolution

Most models support up to 16 million pixels per image. Higher resolutions use more tokens. Token count per image: h x w / (32 x 32) + 2.

Video support

  • Up to 2 hours / 2 GB: qwen3.7-plus, qwen3.6-plus, qwen3.7-flash, qwen3.6-flash, qwen3.5-plus, qwen3.5-flash
  • Up to 1 hour / 2 GB: qwen3-vl-plus, qwen3-vl-flash
  • Up to 1 hour / 2 GB: qwen3.5-omni-plus, qwen3.5-omni-flash (also supports audio input)

Function calling and built-in tools

Allows the model to perform actions based on image or video content.
  • Function calling: Supported by the Qwen3.7, Qwen3.6, Qwen3.5, and Qwen3-VL series.
  • Built-in tools (web search, code execution, no setup required): Available for qwen3.7-plus, qwen3.6-plus, qwen3.7-flash, qwen3.6-flash, qwen3.5-plus, and qwen3.5-flash.

Structured output

Get valid JSON output from visual inputs, such as extracting product details from a photo. Supported by the Qwen3.7, Qwen3.6, Qwen3.5, and Qwen3-VL series in non-thinking mode.

OCR and document extraction

qwen3.5-ocr is optimized for text extraction from documents, tables, exam papers, and handwritten content. For general text extraction from images, use qwen3.7-plus or qwen3.7-flash.

Model ID

Context

Max pixels/image

Max video duration

Max video size

Max images

Max videos

Function calling

Built-in tools

Structured output

qwen3.7-plus

1M

16M

2 hours

2 GB

2048

64

Supported

Supported

Supported

qwen3.7-flash

1M

16M

2 hours

2 GB

256

64

Supported

Supported

Supported

qwen3.5-omni-plus

64k

--

1 hour

2 GB

2,048

512

Supported

--

Supported

Legacy models

Qwen3.7

Model IDInputOutputContextMax outputMax imagesMax videosFunction callingBuilt-in toolsStructured output
qwen3.7-plus
qwen3.7-plus-2026-05-26
Text, images, videoText1M64k204864SupportedSupportedSupported
qwen3.7-flash
qwen3.7-flash-2026-07-15
Text, images, videoText1M64k25664SupportedSupportedSupported

Qwen3.6

Model IDInputOutputContextMax outputMax imagesMax videosFunction callingBuilt-in toolsStructured output
qwen3.6-plus
qwen3.6-plus-2026-04-02
Text, images, videoText1M64k25664SupportedSupportedSupported
qwen3.6-flash
qwen3.6-flash-2026-04-16
Text, images, videoText1M64k25664SupportedSupportedSupported
qwen3.6-35b-a3bText, images, videoText256k64k25664SupportedSupportedSupported

Qwen3.5

Model IDInputOutputContextMax outputMax imagesMax videosFunction callingBuilt-in toolsStructured output
qwen3.5-plus
qwen3.5-plus-2026-02-15
Text, images, videoText1M64k25664SupportedSupportedSupported
qwen3.5-flash
qwen3.5-flash-2026-02-23
Text, images, videoText1M64k25664SupportedSupportedSupported
qwen3.5-397b-a17bText, images, videoText32k8k25664SupportedSupportedSupported
qwen3.5-122b-a10bText, images, videoText32k8k25664SupportedSupportedSupported
qwen3.5-27bText, images, videoText32k8k25664SupportedSupportedSupported
qwen3.5-35b-a3bText, images, videoText32k8k25664SupportedSupportedSupported
These models are no longer recommended. For new projects, use the Qwen3.6 or Qwen3.5 series. For full model specifications, visit the Models page. China (Beijing) | Singapore | U.S. | China (Hong Kong) | Germany (Frankfurt)

Qwen3-VL

  • qwen3-vl-plus and its snapshots
  • qwen3-vl-flash and its snapshots

Qwen-Omni

  • qwen3-omni-flash and its snapshots
  • qwen-omni-turbo and its snapshot

Qwen-OCR

  • qwen-vl-ocr and its snapshots
  • qwen-vl-ocr-latest

QVQ

  • qvq-max
  • qvq-plus

Legacy Qwen-VL

  • qwen-vl-max
  • qwen-vl-plus
Token Plan
Model Playground
Statistics and Monitoring
Support