# Alibaba Cloud Model Studio ## Get Started - [What is Alibaba Cloud Model Studio](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/what-is-model-studio.md): Alibaba Cloud Model Studio is a one-stop model service platform. It provides the full Qwen series and mainstream third-party LLMs through official Qwen APIs and OpenAI-compatible APIs, with multimodal support across text, image, and audio/video. Call models on demand — no infrastructure to manage. - [Make your first API call to Qwen](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/first-api-call-to-qwen.md): Alibaba Cloud Model Studio supports API calls to models through OpenAI-compatible interfaces and the DashScope SDK. - [Recommended models](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/models.md): Alibaba Cloud Model Studio offers Qwen and third-party models for text, image, audio, and video. - [Dynamic rate limiting](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/quota-management.md): Bailian applies dynamic rate limiting to some models. The TPM rate limit value is adjusted monthly based on your Bailian monthly consumption tier, and the actual available TPM is no lower than the rate limit value. - [Rate limiting](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/rate-limit.md): Alibaba Cloud Model Studio applies rate limiting to model calls at the Alibaba Cloud account level, aggregating usage across all RAM users, workspaces, and API keys under the account. Requests are rejected when the limit is exceeded and typically recover automatically within one minute. - [Base URL overview](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/base-url.md): The Base URL is the API endpoint for model calls. Pair it with an API Key from the same billing plan — a mismatch returns a 401 error. API Keys are region-specific. - [Regions and access domains](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/regions.md): Alibaba Cloud Model Studio is available in multiple regions. Each region offers various types of access domains, including workspace-dedicated and DashScope domains, and supports multiple service deployment scopes to meet different access requirements. ## Billing - [Free quota for new users](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/new-free-quota.md): When you first activate Alibaba Cloud Model Studio (Singapore region) , the platform automatically grants you a free quota for various models. - [Model inference pricing](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/model-pricing.md) - [Training and deployment pricing](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/model-training-and-deployment-billing.md): This topic describes the billing rules and pricing for model training and model deployment on Alibaba Cloud Model Studio. - [Savings plans](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/savings-plan-and-resource-package.md): Model Studio offers savings plans and resource plans to help you reduce model costs. - [Billing and cost management](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/bill-query-and-cost-management.md): This topic describes how to query billing details, analyze bills, and stop billing. ## Token Plan - [Token Plan overview](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/token-plan-overview.md): Token Plan is an AI model subscription service offered by Alibaba Cloud Model Studio. It uses Credits as a unified billing unit and supports various AI coding and agent tools. Token Plan offers Personal Edition and Team Edition in two editions, meeting the needs of individual developers to enterprise teams. - [Overview](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/token-plan-personal-overview.md): Token Plan Personal Edition is an AI large model subscription service for individual developers. It uses Credits as a unified billing unit and supports text models, multimodal models, and Harness tools. It is compatible with mainstream AI coding and agent tools. - [Quick start](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/token-plan-personal-quick-start.md): Get started with Token Plan Personal Edition in three steps: choose a plan, obtain an API Key, and configure your AI tools. - [FAQ](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/token-plan-personal-faq.md): Frequently asked questions about Token Plan Personal Edition quotas, purchases, subscriptions, and integration. - [Overview](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/token-plan-team-overview.md): Token Plan Team Edition is an AI large model subscription service offered by Alibaba Cloud Model Studio. It uses Credits as a unified billing unit, supports text generation, image generation, video generation, speech recognition, and real-time voice conversation models, is compatible with mainstream AI coding and agent tools, and provides a team management console, data security guarantees, and stable service operation. - [Quick start](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/token-plan-team-quickstart.md): Complete Token Plan Team Edition subscription and integration in three steps: choose a plan, get your API Key, and configure your AI tool. - [Team management](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/token-plan-team-management.md): Add and manage team members, assign and reclaim seats, and monitor Credits usage in the Token Plan console or management platform. - [FAQ](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/token-plan-team-faq.md): Frequently asked questions about Token Plan Team Edition, covering product selection, Credits billing and quota, model and tool compatibility, usage restrictions, purchasing, renewal, and cancellation. - [Integrate Harness tools](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/token-plan-harness-tool.md): Some Qwen models supported by Token Plan have built-in Harness tools that can extend AI coding tools with capabilities such as web search, code interpreter, and web scraping. - [Integrate multimodal generation models](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/token-plan-multimodal-gen.md): Image generation, video generation, and speech synthesis models in Token Plan must be integrated through each tool's extension mechanism (Skill, Slash Command, or Agent). - [web search](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/web-search-for-coding-plan.md): Add a web search tool to Token Plan supported programming tools so the model can retrieve real-time information. - [Add visual understanding capabilities](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/add-vision-skill.md): Some models supported by Token Plan (qwen3.7-plus, etc.) natively support visual understanding and can process image inputs directly. For text-only models such as glm-5 and MiniMax-M2.5, you can add visual capabilities by configuring a local Skill. - [Coding plan overview](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/coding-plan.md): Coding Plan gives you access to models including Qwen, GLM, Kimi, and MiniMax in popular AI coding tools. For a fixed monthly fee, it offers a cost-effective alternative to pay-as-you-go API billing. - [FAQ](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/coding-plan-faq.md) ## Model Playground - [Text generation](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/text-generation-model.md): Choose the right text generation model for AI agents, chatbots, and document processing. - [Overview](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/text-generation.md): A text generation model generates text from natural language prompts for applications such as chatbots, content creation, document summarization, and code generation. - [Multi-turn conversations](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/multi-round-conversation.md): The Qwen API is stateless. To implement multi-turn conversations, pass conversation history in each request. Use truncation, summarization, or retrieval to manage context and reduce token consumption. - [Streaming output](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/stream.md): In real-time chat or long-text generation applications, long wait times degrade user experience and may trigger server-side timeouts, causing tasks to fail. Streaming output addresses these issues by continuously returning fragments of text as the model generates them. - [Deep thinking](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/deep-thinking.md): Deep thinking models reason before responding, improving accuracy on complex tasks like logical reasoning and math. - [Structured output](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/qwen-structured-output.md): When performing information extraction or structured data generation tasks, a model may return extra text (such as ```json ) that breaks downstream parsing. Enabling structured output ensures the model returns a valid JSON string. The JSON Schema mode also gives you precise control over the output structure and types, eliminating extra validation or retries. - [Partial mode](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/partial-mode.md): For scenarios like code completion and text continuation, you can generate new content starting from an existing text fragment (prefix). Partial Mode ensures the model's output connects seamlessly with your prefix for improved accuracy and control. - [Context Cache](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/context-cache.md): Inference requests for large models often contain overlapping input, such as in a multi-turn conversation or a series of questions about the same book. Context Cache reduces redundant computation by caching the common prefix of these requests. This improves response speed and lowers usage costs without affecting response quality. - [Batch inference](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/batch-inference.md): For inference scenarios that do not require real-time responses, batch inference asynchronously processes large volumes of data requests at 50% of the cost of real-time inference. Its OpenAI-compatible API is ideal for batch jobs such as model evaluation and data labeling. - [Function calling](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/qwen-function-calling.md): Large Language Models (LLMs) cannot access real-time data or external systems. Function Calling enables models to call external tools, such as APIs, databases, and user-defined functions. This allows a model to retrieve information or perform actions beyond its built-in capabilities. - [Web search](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/web-search.md): The training data for large language models has a knowledge cutoff date, preventing them from answering real-time questions. Enabling web search lets the model retrieve real-time data and accurately answer time-sensitive questions, such as stock prices, weather forecasts, and breaking news. - [Web extractor](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/web-extractor.md): LLMs cannot directly access web page data. The web extractor accesses a URL and extracts its content for the model. - [Code Interpreter](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/qwen-code-interpreter.md): Enable the built-in Python Code Interpreter when calling a model. The model writes and runs Python code in a sandbox to solve complex problems such as mathematical calculations and data analytics. - [Text-to-image search](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/web-search-image.md): The text-to-image search tool enables a model to search the Internet for relevant images based on a text description. The model can then describe the image content and perform inference. This is useful for scenarios such as visual Q&A and image recommendations. - [Search by image](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/image-search.md): The search by image tool enables the model to search the Internet for visually similar images based on an input image. The model can then analyze the search results and make inferences. This feature is useful for scenarios such as finding similar products or tracing the origin of visual content. - [Knowledge retrieval](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/file-search.md): Large Language Models (LLMs) cannot answer questions about private data. The knowledge retrieval tool retrieves content from a knowledge base and provides it to the LLM. This enables the model to generate more accurate and relevant answers. - [MCP](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/mcp.md): The Model Context Protocol (MCP) enables large language models to use external tools and data. This topic describes how to connect to MCP using the Responses API. - [PDF understanding](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/pdf-understanding.md): PDF Understanding enables the model to parse and comprehend PDF documents, extracting text and images for analysis. You can pass PDF files via URL or Base64 encoding. - [Long context (Qwen-Long)](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/long-context-qwen-long.md): Qwen-Long handles documents up to 10 million tokens through a file upload and reference mechanism, overcoming standard model context limits. - [Code capabilities (Qwen-Coder)](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/qwen-coder.md): Qwen-Coder is a language model designed for code tasks. You can use the API to generate code, complete code, and call tools to interact with external systems. - [Machine translation (Qwen-MT)](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/machine-translation.md): Qwen-MT is a machine translation model fine-tuned from Qwen3. It supports 92 languages -- including Chinese, English, Japanese, Korean, French, Spanish, German, Thai, Indonesian, Vietnamese, and Arabic -- and offers term intervention, domain prompting, and translation memory to control translation quality. - [Role-playing (Qwen-Character)](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/role-play.md): Qwen's role-playing model enables human-like conversations for virtual social apps, game non-player characters (NPCs), IP replication, and smart hardware such as toys or in-car systems. This model improves character consistency, topic progression, and empathetic listening compared to other Qwen models. - [Data mining (Qwen-Doc-Turbo)](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/data-mining-qwen-doc.md): The data mining model extracts information, moderates content, classifies data, and generates summaries. It outputs structured data (like JSON) quickly and accurately, unlike general-purpose chat models which may return inconsistent formats or extract information incorrectly. - [Deep research (Qwen-Deep-Research)](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/qwen-deep-research.md): Automates complex research through planning, multiple rounds of web searches, and structured report generation. Gathers and synthesizes information without manual effort. - [Mathematical capabilities (Qwen-Math)](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/math-language-model.md): Qwen-Math models provide mathematical reasoning with detailed, verifiable step-by-step solutions. - [Audio understanding (Qwen3-Omni-Captioner)](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/qwen3-omni-captioner.md): Qwen3-Omni-Captioner, an open-source model built on Qwen3-Omni, generates detailed audio descriptions—covering speech, ambient sounds, music, and sound effects—without prompts. It identifies speaker emotions, musical elements (style, instruments), and sensitive information for audio analysis, security audits, intent recognition, and video editing. This model does not support fine-grained acoustic feature analysis such as timbre, pitch, or tone/intonation. - [Visual understanding](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/vision-model.md): Choose the right model for your use case, such as image analysis, video understanding, or OCR. - [Image and video understanding](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/vision.md): Visual understanding models can answer questions based on the images or videos that you provide. They support single or multiple image inputs and are suitable for various tasks, such as image captioning, visual question answering, and object localization. - [Text extraction (Qwen-OCR)](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/qwen-vl-ocr.md): Qwen-OCR is a visual understanding model that extracts text and structured data from images — scanned documents, tables, receipts, and more. It handles multiple languages and supports advanced OCR tasks: information extraction, table parsing, formula recognition, and document parsing. - [Visual reasoning](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/visual-reasoning.md): Visual reasoning models output their thinking process before answering. Use them for complex visual tasks: math problems, chart analysis, or video understanding. - [Image generation and editing](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/image-model.md): Select the right model for text-to-image and image editing. - [Text-to-image](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/text-to-image.md): Generate images from text descriptions with the text-to-image API. This service, provided by Alibaba Cloud Model Studio, features the Wan, Qwen-Image, and Z-Image model families. - [Qwen-Image-Edit](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/qwen-image-edit-guide.md): Qwen-Image-Edit supports multi-image input and output. It can modify text in images, add/delete/move objects, change subject actions, transfer styles, and enhance details. - [Image editing - Wan2.5 to 2.7](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/wan-image-edit.md): Edit images with text instructions using Wan models. Supports multi-image input/output, image fusion, subject preservation, and object detection. - [General image editing - Wan2.1](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/wanx-image-edit.md): Edit images with text prompts using the Wan model: outpainting, watermark removal, style transfer, instruction-based editing, inpainting, and restoration. - [Human instance segmentation](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/image-instance-segmentation.md): Human instance segmentation identifies different human objects in an image and describes each object with a pixel-level mask. - [Video generation and editing](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/video-generate-edit-model.md): Choose the right model for text-to-video, image-to-video, and video editing. - [Wan3.0 Video Generation](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/wan3-video-generation-guide.md): The wan3.0 series is an All-in-One video generation model with comprehensive upgrades in audio-video generation, multi-modal reference, and video editing capabilities. It supports up to 30 seconds per generation, 30fps output frame rate, natively outputs dialogue, BGM, and sound effects, supports up to 20 multi-modal reference materials (images, videos, audio, documents, web pages) per request, and supports first frame/first-last frame control and video editing/extension. - [Text-to-video](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/text-to-video-guide.md): The Wan text-to-video models generate videos from text, images, and audio input. wan3.0-video supports up to 30 seconds, adaptive aspect ratio, smart duration, 480P/720P/1080P resolution, audio toggle, reference audio, and reference files and web links. - [Image-to-video 2.7](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/wan-image-to-video-guide.md): The Wan 2.7 image-to-video model uses multimodal input (text, image, audio, and video) to perform three main tasks: first-frame video generation, first-and-last-frame video generation, and video continuation (continuing from an initial video segment, with or without a final frame) . - [Image-to-video: first and last frames](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/image-to-video-first-and-last-frames-guide.md): The Wan image-to-video model generates smooth videos from a first-frame image, a last-frame image, and an optional text prompt. - [image-to-video - based on first frame](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/image-to-video-guide.md): The Wan Image-to-Video model accepts multimodal input (text, image, or audio) and generates videos up to 15 seconds long at 1080P resolution. - [Reference-to-video](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/video-to-video-guide.md): Wan-R2V accepts multimodal input (text, image, video, and audio) to generate performance videos. Use prompts to cast people or objects as the main characters. - [Video Editing 2.7](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/wan-video-editing-guide.md): The Wanxiang video editing model supports editing operations on input videos, such as adding, deleting, or modifying content, replacing backgrounds, converting styles, and replicating actions, effects, or camera movements. It offers the following two methods: - [Video editing 2.1](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/wan-vace-guide.md): The Wanxiang universal video editing model supports multimodal inputs (text, image, and video) and provides five core capabilities: multi-image reference, video repainting, local editing, video extension, and video outpainting . - [Speech synthesis](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/tts-model.md): Choose the right model for speech synthesis, voice cloning, and voice design. - [Real-time speech synthesis](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/realtime-tts-user-guide.md): Stream text-to-speech conversion with low first-packet latency. Real-time speech synthesis supports streaming input and output, voice cloning, voice design, and fine-grained audio controls for voice assistants, audiobooks, and intelligent customer service. - [Non-real-time speech synthesis](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/non-realtime-tts-user-guide.md) - [Voice cloning](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/voice-cloning-user-guide.md): Voice cloning requires only a 10–20 second audio sample to generate a highly similar custom voice without model training. - [Voice Design](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/voice-design-user-guide.md) - [SSML](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/ssml-latex-user-guide.md): Use SSML (Speech Synthesis Markup Language) to fine-tune speech characteristics such as speed, pauses, and pronunciation. - [Qwen-Audio-TTS voice list](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/qwen-audio-tts-voice-list.md) - [CosyVoice voice list](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/cosyvoice-voice-list.md) - [Qwen-TTS voice list](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/qwen-tts-voice-list.md) - [Music generation](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/fun-music.md): Fun-Music generates complete songs with male or female vocals in Chinese or English from a text prompt describing the music style and scene, or from custom lyrics. - [Speech-to-text](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/asr-model.md): Choose the right model for real-time or file-based speech recognition. - [Real-time speech recognition - Qwen](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/real-time-speech-recognition-user-guide.md): The real-time speech recognition service receives an audio stream and transcribes it into punctuated text in real time. Use it for live captioning, online meetings, voice chat, smart assistants, and similar scenarios. - [Non-real-time speech recognition](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/non-realtime-speech-recognition-user-guide.md): Non-real-time speech recognition models convert recorded audio into text. They support multilingual recognition, singing recognition, noise rejection, and speaker diarization, which makes them suitable for meeting transcription, call analysis, subtitle generation, and similar scenarios. - [Improve recognition accuracy](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/improve-asr-accuracy.md): Model Studio speech recognition offers two ways to improve the recognition accuracy of specialized terms, product names, and other domain-specific vocabulary: custom hotwords and context enhancement. This topic describes the scope and usage of each approach. - [Speech-to-speech](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/s2s-model.md): Choose a model for voice conversation, speech translation, or simultaneous interpretation. - [Qwen-Audio real-time voice model](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/qwen-audio-realtime-user-guides.md): Qwen-Audio is an end-to-end real-time voice interaction model for low-latency voice conversations. Use cases include voice assistants, intelligent customer service, and AI companions. - [Real-time audio and video translation - Qwen](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/qwen3-5-livetranslate-flash-realtime.md): qwen3.5-livetranslate-flash-realtime is a vision-enhanced real-time translation model supporting 60 languages (29 with audio + text, 31 text-only). It processes audio and image input from video streams or local files, uses visual context to improve accuracy, and outputs translated text and audio in real time. - [Audio and Video File Translation – Qwen](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/qwen3-livetranslate-flash.md): qwen3-livetranslate-flash translates audio and video files across 18 languages. It accepts audio or video input and returns translated text, synthesized audio, or both via a streaming API. For video input, visual context improves translation accuracy (e.g., distinguishing "medical mask" vs. "masquerade mask" based on video frames). - [Omni-modal](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/omni.md): Choose a model for voice conversation, audio and video analysis, content moderation, voice translation, and other omni-modal tasks. - [Qwen-Omni-Realtime](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/realtime.md): Qwen-Omni-Realtime processes streaming audio and image inputs (including video frames) and generates text and audio responses in real time. - [Qwen-Omni](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/qwen-omni.md): The Qwen-Omni model accepts multimodal input and generates text or speech responses. It produces human-like voices and supports speech output in multiple languages and dialects. Use cases include content moderation, text creation, visual recognition, and audio-video interaction. - [Voice list](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/omni-voice-list.md): Voices supported by Qwen-Omni (non-real-time) and Qwen-Omni-Realtime models, with their corresponding voice parameter values. - [Embedding and rerank](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/embedding-rerank-model.md): Find the right models for semantic search, Retrieval-Augmented Generation (RAG), cross-modal matching, and reranking. - [Embedding](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/embedding.md): Embedding models convert data such as text, images, and videos into vectors for downstream tasks, including semantic search, recommendation, clustering, classification, and anomaly detection. - [Rerank](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/rerank.md): A reranking model rescores retrieved documents and promotes the most relevant results, improving search precision when initial retrieval prioritizes speed over accuracy. ## Clients and Developer Tools - [OpenClaw](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/openclaw.md): OpenClaw is an open source personal AI assistant platform that lets you interact with AI through various messaging channels. You can configure it to access AI models from Alibaba Cloud Model Studio. It supports four access methods: pay-as-you-go, Coding Plan, Token Plan Personal Edition, and Token Plan Team Edition. - [Hermes Agent](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/hermes-agent.md): Hermes Agent is a terminal AI coding tool that connects to Alibaba Cloud Model Studio through Pay-as-you-go, Coding Plan, Token Plan Personal Edition, or Token Plan Team Edition. - [Claude Code](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/claude-code.md): Claude Code is a command-line AI coding assistant developed by Anthropic. Connect it to Alibaba Cloud Model Studio using Pay-as-you-go, Coding Plan, Token Plan Personal Edition, or Token Plan Team Edition. - [OpenCode](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/opencode.md): OpenCode is a terminal-based AI coding assistant. You can connect it to Alibaba Cloud Model Studio using a pay-as-you-go plan, Coding Plan, Token Plan Personal Edition, or Token Plan Team Edition. - [Cursor](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/cursor.md): Cursor is an AI coding IDE. You can connect it to Alibaba Cloud Model Studio using Pay-as-you-go, Coding Plan, Token Plan Personal Edition, or Token Plan (Team Edition). - [Codex](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/codex.md): Codex is OpenAI's terminal AI coding assistant. Connect it to Alibaba Cloud Model Studio via Token Plan Personal Edition, Token Plan Team Edition, Coding Plan, or pay-as-you-go. - [Qwen Code](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/qwen-code.md): Qwen Code is a terminal-based AI coding tool that can be connected to Alibaba Cloud Model Studio via pay-as-you-go, Coding Plan, Token Plan (Personal Edition), or Token Plan (Team Edition). - [DeepSeek Harness](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/deepseek-harness.md): DeepSeek Harness is an open-source AI Agent framework by DeepSeek, supporting both Web UI and CLI modes. Connect it to Alibaba Cloud Model Studio using Pay-as-you-go, Coding Plan, Token Plan Personal Edition, or Token Plan Team Edition. - [QwenPaw](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/qwenpaw.md): QwenPaw (formerly CoPaw) is an open-source personal AI assistant from the AgentScope team. It supports local and cloud deployment, and integrates with Alibaba Cloud Model Studio via Token Plan Personal Edition, Token Plan Team Edition, Coding Plan, or Pay-as-you-go. - [Cherry Studio](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/cherry-studio.md): Cherry Studio is an open-source AI desktop client. You can connect it to Model Studio using Token Plan Personal Edition, Token Plan Team Edition, Coding Plan, or Pay-as-you-go. - [Chatbox](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/chatbox.md): Chatbox is a cross-platform AI client application. Connect it to Alibaba Cloud Model Studio via Token Plan Personal Edition, Token Plan Team Edition, Coding Plan, or pay-as-you-go. - [Cline](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/cline.md): Cline is a VS Code extension for AI-assisted coding. You can connect it to Alibaba Cloud Model Studio using Token Plan Personal Edition, Token Plan Team Edition, Coding Plan, or Pay-as-you-go. - [Qoder](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/qoder-agent.md): Qoder is an agentic coding platform for software development that supports a desktop IDE, CLI, and JetBrains plugin, and can connect to Alibaba Cloud Model Studio via pay-as-you-go, Coding Plan, Token Plan Personal Edition, or Token Plan Team Edition. - [Qoder CN (formerly Lingma)](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/lingma-agent.md): Qoder CN (formerly Lingma) is Alibaba Cloud's intelligent coding assistant that provides a standalone IDE, and can connect to Alibaba Cloud Model Studio via Token Plan Personal Edition, Token Plan Team Edition, Coding Plan, or pay-as-you-go. - [Kilo CLI](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/kilo-cli.md): Kilo CLI is the command-line client for Kilo Code. It connects to Alibaba Cloud Model Studio via pay-as-you-go, Coding Plan, Token Plan Personal Edition, or Token Plan Team Edition. - [Use Postman or cURL to call image and video generation APIs](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/first-call-to-image-and-video-api.md): This article describes how to use Postman and cURL to call image or video generation APIs on Alibaba Cloud Model Studio (Bailian). This guide uses a text-to-image example to demonstrate the complete process, from creating a task to retrieving the result. - [Dify](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/dify.md): Dify is an open-source platform for building AI applications. You can build applications using model APIs from Alibaba Cloud Model Studio. - [More tools](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/more-tools.md): Model Studio supports any third-party programming tool compatible with OpenAI or Anthropic API protocols that allows custom endpoints. Connect via Pay-as-you-go, Coding Plan, Token Plan Personal Edition, or Token Plan (Team Edition). ## Model Inference - [Fast mode](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/fast-mode.md): Fast mode provides higher TPS for latency-sensitive scenarios. - [TPM reservation](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/tpm-reservation.md): A TPM reservation locks dedicated inference capacity for a specified model, ensuring that your services are not affected by public rate limits during peak business hours. This topic describes how to create, integrate, and manage TPM reservations. ## Fine-tuning - [Introduction to model fine-tuning](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/model-training-overview.md): You can use model fine-tuning in Alibaba Cloud Model Studio if a model's performance does not meet your expectations after you have tried optimization methods such as prompt engineering and plugin calls. As a core strategy for improving model performance, model fine-tuning can significantly enhance a model's capabilities in specific industries or business scenarios, align its outputs with human preferences, and reduce output latency. Model fine-tuning includes three training methods: supervised fine-tuning (SFT), continual pre-training (CPT), and direct preference optimization (DPO). - [Fine-tuning data upload rules](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/text-generation-tuning-data-upload-rules.md): Describes the format specifications, packaging requirements, size and quantity limits, and API upload quotas for text generation model fine-tuning data, helping users construct and upload compliant SFT/DPO/CPT training data by training method. - [Fine-tune models in the console](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/model-training-on-console.md): This topic describes how to run model fine-tuning tasks in the console and helps you choose the right fine-tuning method and parameters. Model fine-tuning includes three training methods: Supervised Fine-Tuning (SFT), Continual Pre-Training (CPT), and Direct Preference Optimization (DPO). - [Fine-tune with the API or CLI](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/fine-tuning-api-guide.md): Tune Qwen models in Model Studio through the API (HTTP) or CLI (shell). Three tuning methods are supported: supervised fine-tuning (SFT), continual pre-training (CPT), and direct preference optimization (DPO). - [Fine-tune image generation models](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/wan-image-generation-finetune-guide.md): When using Wan for image generation , if Text-to-video/image-to-video prompt guide cannot meet your customization needs for specific styles, IP characters, or visual effects , use model fine-tuning . - [Fine-tuning video generation models](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/wan-video-generation-finetune-guide.md): When using for image-to-video , if prompt optimization or calling the official video effects does not meet your custom requirements for specific actions, effects, or styles , use model fine-tuning . ## Deployment - [Model deployment](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/model-deployment-introduction.md): you can obtain an independent, resource-dedicated inference service through deployment to meet your business needs for different performance levels such as high concurrency and low latency. - [Provisioned Throughput Long Input and Cache](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/ptu-long-input-and-cache.md): This topic describes the long-input and prefix cache capabilities of PTU (Provisioned Throughput) deployments, including quota consumption rules, how to use the Provisioned Throughput Quota Calculator , and API response field descriptions. - [Model Import](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/model-import.md): My Models lets you manage models you created or imported. Use this page to import a locally trained LoRA model from Alibaba Cloud Object Storage Service (OSS) into Alibaba Cloud Model Studio. - [Deploy a model by using an API or CLI](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/model-deployment-quick-start.md): This topic shows you how to deploy a Qwen model on Alibaba Cloud Model Studio by using API calls. - [DTU Model Deployment](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/dtu-model-deployment.md): Dedicated Throughput Unit (DTU) provides performance and throughput guarantees for specified models through dedicated deployment, billed by input/output TPM (Tokens Per Minute). This document introduces DTU features, billing, usage flow, and supported models. ## Evaluation - [Model Evaluation](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/model-evaluation-intl.md): Model Evaluation is a model quality evaluation tool provided by the Alibaba Cloud Model Studio platform. It supports quantitative evaluation of large language model performance through custom evaluation dimensions, helping you complete model selection, tuning verification, and capability comparison. ## Statistics and Monitoring - [Model usage](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/model-usage-statistics.md): View usage for Alibaba Cloud Model Studio models. - [Model monitoring](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/model-telemetry.md): Model Studio provides model monitoring, alerting, and logging so you can monitor model calls in real time, detect anomalies promptly, and troubleshoot issues. ## Security and Compliance - [Permission management](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/permission-management-overview.md): Alibaba Cloud Model Studio provides granular access control at the console and model levels to support complex organizational structures spanning multiple regions. - [Access Model Studio APIs over a private network](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/access-model-studio-through-privatelink.md): To call Model Studio APIs from a VPC without routing traffic over the public internet, create a PrivateLink endpoint. - [Security certifications and privacy](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/privacy-notice.md) - [Alibaba Cloud Model Studio - Training data summary](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/qwen-and-wan-training-data-disclosure.md) ## Best Practices - [Text-to-text prompt guide](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/prompt-engineering-guide.md): A prompt is the text you input to a large language model (LLM), used to explicitly tell the model what problem you want to solve or what task you want to complete. It is also the foundation for the LLM to understand user needs and generate relevant, accurate answers or content. To help you use LLMs more efficiently, this tutorial provides a series of practical techniques to help you design and optimize prompts. - [Text-to-image prompt guide](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/text-to-image-prompt.md): Write effective prompts to generate high-quality images with Wan - text-to-image V2 . This guide covers prompt structure, visual vocabulary, and practical examples. - [Text-to-video/image-to-video prompt guide](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/text-to-video-prompt.md): Use structured prompt formulas to control content, motion, camera work, audio, and visual style in AI-generated videos. - [Best practices for handling rate limiting](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/rate-limiting-best-practices.md): Model Studio APIs limit request volume, token usage, and growth rate. Apply these strategies to maximize throughput and maintain availability. - [Explicit Cache Best Practices](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/explicit-cache-guide.md): This topic describes how to use explicit cache and its best practices. By adding cache markers to your requests, explicit cache guarantees deterministic cache hits for identical input content, significantly reducing cost and latency. - [DeepSeek](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/deepseek-api.md): This topic describes how to call DeepSeek series models on the Alibaba Cloud Model Studio platform using an OpenAI compatible interface or the DashScope SDK. - [Kimi](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/kimi-api.md): This document describes how to call the Kimi model inference service deployed on Alibaba Cloud Model Studio. - [GLM](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/glm.md): This topic describes how to use APIs to call GLM series models on the Alibaba Cloud Model Studio platform. - [GLM-ZHIPU](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/glm-zhipu.md): This document describes how to call the Z.AI model inference service on Alibaba Cloud Model Studio. - [MiniMax](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/minimax-api.md): Call MiniMax models on Alibaba Cloud Model Studio. - [Using WebRTC with qwen3.5-omni-plus-realtime for real-time calls](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/best-practice-webrtc-omni-realtime.md): Describes how to connect to the Realtime API in a browser using WebRTC and JavaScript to enable real-time audio/video calls with the qwen3.5-omni-plus-realtime model. - [Using AOQ with qwen3.5-omni-plus-realtime for real-time calls](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/best-practice-aoq-omni-realtime.md): Describes how to integrate the AOQ Client SDK on Android, iOS, and HarmonyOS to implement audio/video calls using AOQ and qwen3.5-omni-plus-realtime. - [Build push-to-talk voice conversations with qwen3.5-omni-plus-realtime over AOQ](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/use-aoq-to-access-qwen3-5-omni-plus-realtime-to-realize-key-voice-dialogue.md): Use AOQ to connect to qwen3.5-omni-plus-realtime and let the client control turn boundaries for push-to-talk conversations and optional image questions. The client code uses iOS Swift. - [Build real-time voice conversations with qwen-audio-3.0-realtime-plus over AOQ](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/real-time-voice-conversation-using-aoq-access-qwen-audio-3-0-realtime-plus.md): Use AOQ to connect to qwen-audio-3.0-realtime-plus and use server-side VAD to build low-latency real-time voice conversations. The client code uses Android Java. - [Synthesize speech with qwen-audio-3.0-tts-flash over AOQ](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/speech-synthesis-using-aoq-access-qwen-audio-3-0-tts-flash.md): Use AOQ to connect to qwen-audio-3.0-tts-flash, send text in segments, and play synthesized speech in real time. The client code uses Android Java. - [Transcribe real-time speech with fun-asr-realtime over AOQ](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/real-time-speech-recognition-using-aoq-access-fun-asr-realtime.md): Use AOQ to connect to fun-asr-realtime, stream microphone audio, and receive transcription results in real time. The client code uses Android Java, and other AOQ-supported platforms provide the same interfaces. ## Support - [qwen3.8-max](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/qwen3-8-max.md): A 2.4-trillion-parameter MoE flagship with a major leap in coding and office productivity, able to autonomously code for over ten days to deliver complete projects. It handles hundreds of professional tasks across law, finance, design, and more, delivering production-grade results end-to-end in a single conversation. Native visual understanding runs through the entire planning, execution, and verification pipeline, enabling deep semantic parsing of ultra-long documents and long videos. It plans autonomously and iterates in closed loops during long-horizon tasks, continuously improving. - [qwen3.7-max](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/qwen3-7-max.md): The Max model, the largest and most capable in the Qwen3.7 series, currently offers a pure‑text‑only interface for public experimentation. Qwen3.7 is a next‑generation flagship model designed for the agent‑centric era, with its core strengths lying in the breadth and depth of its agent‑level capabilities: it excels at programming, office and productivity tasks, and long‑term autonomous execution.This model version is functionally equivalent to the snapshot model qwen3.7-max-2026-05-20. - [qwen3.7-max-us](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/qwen3-7-max-us.md): The Max model, the largest and most capable in the Qwen3.7 series, currently offers a pure‑text‑only interface for public experimentation. Qwen3.7 is a next‑generation flagship model designed for the agent‑centric era, with its core strengths lying in the breadth and depth of its agent‑level capabilities: it excels at programming, office and productivity tasks, and long‑term autonomous execution. - [qwen3.6-max-preview](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/qwen3-6-max.md): The Max model, the largest and most capable variant in the Qwen3.6 series, is now available in a preview version. At present, only its plain-text capabilities are open for experimentation. Compared with the previously released Qwen3-Max and Qwen3.6-Plus, this model features enhanced vibe coding abilities, more efficient coding agent execution, and significantly improved front-end development skills. Additionally, its long-tail knowledge retention has been further upgraded. - [qwen3-max](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/model-qwen3-max.md): The Qwen 3 series Max model has undergone specialized upgrades in agent programming and tool invocation compared to the preview version. The officially released model this time has achieved state-of-the-art (SOTA) performance in its field and is better suited to meet the demands of agents operating in more complex scenarios.This model version is functionally equivalent to the snapshot model qwen3-max-2026-01-23. - [qwen-max](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/qwen-max.md): Qwen-Max supports a parameter scale of hundreds of billions and multiple input languages such as Chinese and English. Qwen-Max is updated in a rolling manner. Compared to previous versions, it shows significant improvements in both Chinese and English code generation, logical reasoning, and multilingual abilities. The response style has been greatly adjusted to align with human preferences, with noticeable enhancements in the level of detail and clarity of responses. Specialized improvements have been made in creative writing, adherence to JSON formatting, and role-playing abilities.This model version is functionally equivalent to the snapshot model qwen-max-2025-01-25. - [qwen3.7-plus](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/qwen3-7-plus.md): Among the Qwen3.7 series, the cost-effective Plus model builds on its robust text capabilities while delivering a comprehensive upgrade to its vision‑language abilities, all while preserving its full‑stack agent‑level intelligence for coding, tool use, and productivity workflows. Its key distinguishing feature is multi‑modal interactive hybrid agent capabilities, enabling it to perceive real‑world scenes, read screens and interact with GUIs, generate code based on visual references, and perform end‑to‑end navigation within mobile apps.This model version is functionally equivalent to the snapshot model qwen3.7-plus-2026-05-26. - [qwen3.7-plus-us](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/qwen3-7-plus-us.md): Among the Qwen3.7 series, the cost-effective Plus model builds on its robust text capabilities while delivering a comprehensive upgrade to its vision‑language abilities, all while preserving its full‑stack agent‑level intelligence for coding, tool use, and productivity workflows. Its key distinguishing feature is multi‑modal interactive hybrid agent capabilities, enabling it to perceive real‑world scenes, read screens and interact with GUIs, generate code based on visual references, and perform end‑to‑end navigation within mobile apps. - [qwen3.6-plus](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/qwen3-6-plus.md): The Qwen3.6 native vision-language Plus series models demonstrate exceptional performance on par with the current state-of-the-art models, with a significant improvement in overall results compared to the 3.5 series. The models have been markedly enhanced in code-related capabilities such as agentic coding, front-end programming, and Vibe coding, as well as in multi-modal general object recognition, OCR, and object localization.This model version is functionally equivalent to the snapshot model qwen3.6-plus-2026-04-02. - [qwen3.5-plus](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/qwen3-5-plus.md): The Qwen3.5 native vision-language series Plus models are built on a hybrid architecture that integrates linear attention mechanisms with sparse mixture-of-experts models, achieving higher inference efficiency. In a variety of task evaluations, the 3.5 series consistently demonstrates performance on par with state-of-the-art leading models. Compared to the 3 series, these models show a leap forward in both pure-text and multimodal capabilities. This model version is functionally equivalent to the snapshot model qwen3.5-plus-2026-02-15. - [qwen-plus](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/qwen-plus.md): Qwen-Plus is an enhanced version of the Qwen ultra-large language model that supports multiple input languages such as Chinese and English. Compared to previous versions, it shows significant improvements in both Chinese and English code generation, logical reasoning, and multilingual abilities. The response style has been greatly adjusted to align with human preferences, with noticeable enhancements in the level of detail and clarity of responses. Specialized improvements have been made in creative writing, adherence to JSON formatting, and role-playing abilities.This model version is functionally equivalent to the snapshot model qwen-plus-2025-07-28. - [qwen-plus-us](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/qwen-plus-us.md): Qwen-Plus is an enhanced version of the Qwen ultra-large language model that supports multiple input languages such as Chinese and English. Compared to previous versions, it shows significant improvements in both Chinese and English code generation, logical reasoning, and multilingual abilities. The response style has been greatly adjusted to align with human preferences, with noticeable enhancements in the level of detail and clarity of responses. Specialized improvements have been made in creative writing, adherence to JSON formatting, and role-playing abilities.This model version is functionally equivalent to the snapshot model qwen-plus-2025-12-01-us. - [qwen3.8-flash](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/qwen3-8-flash.md): Qwen3.8-Flash is the latest multimodal model from the Qwen family, combining powerful reasoning and generation with remarkable speed. It natively supports a million-token context window, allowing it to process lengthy documents, entire codebases, and complex conversations in a single pass. It shines in coding assistance, agentic workflows, and visual understanding — whether it's fixing code autonomously, operating desktop applications, or analyzing charts and long videos. Fully compatible with both OpenAI and Anthropic API protocols, it integrates seamlessly with popular developer tools like Claude Code and Codex, making it easy to build high-concurrency applications and intelligent workflows. With strong performance and highly competitive inference costs, Qwen3.8-Flash is an ideal choice for developers and businesses seeking the best of both worlds in AI applications. - [qwen3.7-flash](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/qwen3-7-flash.md): The Qwen3.7 native vision-language series Flash models comprehensively enhance multimodal understanding and Agent execution capabilities compared to 3.6-Flash. Key improvements include strengthened foundational multimodal abilities, enhanced object recognition, improved real-world perception and spatial intelligence. Multimodal Agent scenarios such as Search Agent and CI Agent have seen significant upgrades, with more stable end-to-end task execution. Multimodal coding capabilities are optimized, delivering a smoother vibe coding experience. - [qwen3.6-flash](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/qwen3-6-flash.md): The Qwen3.6 native vision-language Flash model series delivers a significant performance boost over the 3.5-Flash version. This model particularly excels in agentic coding capabilities, substantially outperforming its predecessor on multiple code-agent benchmarks, as well as in mathematical and code reasoning. In terms of vision, it features markedly improved spatial intelligence, with especially notable enhancements in object localization and object detection. - [qwen3.6-flash-us](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/qwen3-6-flash-us.md): The Qwen3.6 native vision-language Flash model series delivers a significant performance boost over the 3.5-Flash version. This model particularly excels in agentic coding capabilities, substantially outperforming its predecessor on multiple code-agent benchmarks, as well as in mathematical and code reasoning. In terms of vision, it features markedly improved spatial intelligence, with especially notable enhancements in object localization and object detection. - [qwen3.5-flash](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/qwen3-5-flash.md): The Qwen3.5 native vision-language Flash models are built on a hybrid architecture that integrates a linear attention mechanism with a sparse mixture-of-experts model, achieving higher inference efficiency. Compared to the 3 series, these models deliver a leap forward in performance for both pure text and multimodal tasks, offering fast response times while balancing inference speed and overall performance.This model version is functionally equivalent to the snapshot model qwen3.5-flash-2026-02-23. - [qwen-flash](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/qwen-flash.md): The Qwen3 Flash model offers a powerful fusion of thinking and non-thinking modes with dynamic in-conversation switching, excelling in complex reasoning while showing significant gains in instruction following and text comprehension. It supports a 1M context length and is billed on a tiered model corresponding to context usage. This model version is functionally equivalent to the snapshot model qwen-flash-2025-07-28. - [qwen-flash-us](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/qwen-flash-us.md): The Qwen3 Flash model offers a powerful fusion of thinking and non-thinking modes with dynamic in-conversation switching, excelling in complex reasoning while showing significant gains in instruction following and text comprehension. It supports a 1M context length and is billed on a tiered model corresponding to context usage.This model version is functionally equivalent to the snapshot model qwen-flash-2025-07-28-us. - [qwen-turbo](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/qwen-turbo.md): The Turbo model of the Qwen3 series. It effectively integrates thinking mode and non-thinking mode, allowing for mode switching during conversations. Its reasoning capabilities rival those of QwQ-32B with a smaller parameter size, while its general capabilities significantly surpass those of Qwen2.5-Turbo, achieving the SOTA level in the same scale within the industry.This model version is functionally equivalent to the snapshot model qwen-turbo-2025-04-28. - [qwq-plus](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/qwq-plus.md): The enhanced version of the Qwen QwQ reasoning model, trained on the Qwen2.5 model, has significantly improved its reasoning capabilities through reinforcement learning. The model's core metrics in mathematics and coding (e.g., AIME 24/25, LiveCodeBench) as well as some general metrics (e.g., IFEval, LiveBench) have reached the level of the full version of DeepSeek-R1. - [qwen-plus-character](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/qwen-plus-character.md): The role-playing model of the Qwen series. This is a dynamically updated version, and notifications will be provided in advance for any model updates. It is suitable for anthropomorphic role-playing and has optimized capabilities in following predefined character instructions, advancing conversations, and demonstrating active listening and empathy. Additionally, it supports the deep restoration of personalized characters. - [qwen-plus-character-ja](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/qwen-plus-character-ja.md): The Qwen Role-Playing Model Series is specifically optimized for Japanese anthropomorphic interaction scenarios. It demonstrates advanced capabilities in character consistency maintenance, context-aware dialogue progression, and empathetic engagement, enabling precise personalized character embodiment. This version significantly enhances Japanese linguistic localization (including dialects and honorifics), human-like role-playing authenticity, narrative coherence control, and scenario-based cognitive intelligence. - [qwen-flash-character](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/qwen-flash-character.md): The Qwen Role-Playing Model Series is specifically optimized for muti-language anthropomorphic interaction scenarios. It demonstrates advanced capabilities in character consistency maintenance, context-aware dialogue progression, and empathetic engagement, enabling precise personalized character embodiment. This version significantly enhances Japanese linguistic localization (including dialects and honorifics), human-like role-playing authenticity, narrative coherence control, and scenario-based cognitive intelligence. - [qwen-long](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/qwen-long.md): Qwen-Long is a large language model designed for ultra-long context processing, supporting Chinese, English, and other languages. It handles up to 10 million tokens (approximately 15 million characters or 15,000 document pages) in dialogues. Integrated with document services, it supports parsing and dialogue for text files (TXT, DOCX, PDF, XLSX, EPUB, MOBI, MD, CSV) and image files (BMP, PNG, JPG/JPEG, GIF, PDF scans). Notes: HTTP requests support up to 1M tokens; file submission is recommended for longer content. - [qwen-math-plus](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/qwen-math-plus.md): Qwen-Math-Plus is a powerful math problem-solving model excelling in Chinese and English mathematical tasks, including equations, calculations, and proofs. - [qwen-math-turbo](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/qwen-math-turbo.md): Qwen-Math-Turbo is a specialized language model for math problem-solving, featuring high inference speed and low cost. - [qwen3-coder-flash](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/qwen3-coder-flash.md): Based on Qwen3, this code generation model inherits the coding agent capabilities of Qwen3-Coder-Plus and supports multi-turn tool interaction. It features focused optimizations on repository-level understanding and enhanced tool-calling stability. This model version is functionally equivalent to the snapshot model qwen3-coder-flash-2025-07-28. - [qwen3-coder-next](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/qwen3-coder-next.md): The new-generation code generation model in the Qwen3 series delivers performance close to that of Qwen3-Coder-Plus while offering even better capabilities. The model has been optimized with a focus on repository-level understanding, supports multi-turn tool interactions, and enhances its compatibility with agentic coding tools. - [qwen3-coder-plus](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/qwen3-coder-plus.md): Qwen3-based code generation model with strong coding agent power, excels at tool calling and environment interaction, capable of autonomous programming with outstanding code capability while maintaining general ability.This model version is functionally equivalent to the snapshot model qwen3-coder-plus-2025-07-22. - [qwen-coder-plus](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/qwen-coder-plus.md): Qwen-Coder-Plus is a specialized language model for programming and code generation, delivering excellent performance and outstanding results. - [qwen-coder-turbo](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/qwen-coder-turbo.md): Qwen-Coder-Turbo is a specialized language model for programming and code generation, featuring high inference speed and low cost. - [qwen-mt-plus](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/qwen-mt-plus.md): Qwen-MT-Plus, the flagship translation model from our Qwen series, is now fully upgraded with the Qwen3 architecture. It supports 92 languages and delivers exceptionally accurate and natural-sounding translations. Its advanced capabilities in contextual understanding, terminology control, and format preservation make it a superior choice over traditional models, especially for specialized domains. - [qwen-mt-flash](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/qwen-mt-flash.md): Qwen-MT-Flash, a large language model from the Qwen series, has been fully upgraded with the Qwen 3 architecture for significantly enhanced performance and translation quality. It provides rapid, cost-effective translation across 92 languages, while supporting advanced features such as terminology intervention, format preservation, and domain-specific adaptation. It is the ideal choice for applications requiring a powerful balance of speed, quality, and cost. - [qwen-mt-lite](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/qwen-mt-lite.md): Qwen-MT-Lite is a large language model of the Qwen model series that specializes in multi-lingual translation. It provides high-quality and rapid translation services across 32 languages at a cost-effective price. It offers features such as terminology intervention, format preservation, and domain-specific translation to cater to the diverse needs of various applications, ensuring both efficiency and performance. - [qwen-mt-lite-us](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/qwen-mt-lite-us.md): Qwen-MT-Lite is a large language model of the Qwen model series that specializes in multi-lingual translation. It provides high-quality and rapid translation services across 32 languages at a cost-effective price. It offers features such as terminology intervention, format preservation, and domain-specific translation to cater to the diverse needs of various applications, ensuring both efficiency and performance. - [qwen-mt-turbo](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/qwen-mt-turbo.md): Qwen-MT-Turbo is a large language model within the Qwen model series that specializes in multi-lingual translation. It provides high-quality and rapid translation services across 92 languages at a cost-effective price point. It also offers features such as terminology intervention, format preservation, and domain-specific translation to cater to the diverse needs of various applications, ensuring both efficiency and performance. - [qwen-doc-turbo](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/qwen-doc-turbo.md): Qwen-Doc-Turbo rapidly extracts precise information from documents, supporting tagging, classification, content moderation, and summarization. - [qwen-deep-research](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/model-qwen-deep-research.md): Qwen-Deep-Research is an advanced agent system for complex research tasks, equipped with multi-step reasoning and global planning capabilities. It leverages tools like internet search to perform detailed task decomposition, reasoning, and analysis, ultimately generating traceable, logically rigorous research reports. - [tongyi-intent-detect-v3](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/tongyi-intent-detect-v3.md): Intent Detection and Slot Filling model for dialogue systems, enabling joint prediction of API-based intents and slot parameters in a single output, returning standardized JSON results with multiple API commands and filled slots. - [qwen3-coder-30b-a3b-instruct](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/qwen3-coder-30b-a3b-instruct.md): Qwen3-based code generation model that inherits the coding agent ability of Qwen3-Coder-480B-A35B-Instruct; code capability reaches SOTA at the same scale. - [qwen3-coder-480b-a35b-instruct](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/qwen3-coder-480b-a35b-instruct.md): Qwen3-based code generation model with strong coding agent power; code capability reaches open-source SOTA. - [qwen3.8-2.4t-a95b](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/qwen3-8-2-4t-a95b.md): Qwen3.8-2.4T-A95B is the open-source release of Qwen's latest flagship, launched August 2026. Its sparse MoE architecture holds 2.4T total parameters with ~95B activated per step, paired with hybrid attention and a 1M token context window. Key benchmarks: GPQA Diamond 92.6, PaperBench 93.0, OSWorld 86.1, BabyVision 82.0. Ranked 4th on CodeArena. - [qwen3.8-27b](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/qwen3-8-27b.md): The Qwen3.8 27B native vision-language dense model builds upon the 3.6-27B version, with key improvements in coding and office productivity capabilities across both text and visual modalities. It enables more reliable end-to-end completion of complex tasks, delivering consistently trustworthy results. - [qwen3.6-27b](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/qwen3-6-27b.md): The Qwen3.6 27B native vision-language dense model builds upon the 3.5-27B version, with key improvements in agentic coding capabilities and enhanced STEM reasoning and inference skills. In the vision modality, it demonstrates significant advances in spatial intelligence, object localization, and detection, while video understanding, document OCR, and visual agent capabilities continue to improve steadily. - [qwen3.6-35b-a3b](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/qwen3-6-35b-a3b.md): The Qwen3.6 35B-A3B native vision-language model is built on a hybrid architecture that integrates linear attention mechanisms with a sparse mixture-of-experts framework, achieving higher inference efficiency. Compared with the 3.5-35B-A3B, this model demonstrates significantly improved agentic coding capabilities, mathematical and code reasoning abilities, spatial intelligence, as well as object localization and object detection performance. - [qwen3.5-122b-a10b](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/qwen3-5-122b-a10b.md): The Qwen3.5 122B-A10B native vision-language model is built on a hybrid architecture that integrates a linear attention mechanism with a sparse mixture-of-experts model, achieving higher inference efficiency. In terms of overall performance, this model is second only to Qwen3.5-397B-A17B. Its text capabilities significantly outperform those of Qwen3-235B-2507, and its visual capabilities surpass those of Qwen3-VL-235B. - [qwen3.5-27b](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/qwen3-5-27b.md): The Qwen3.5 27B native vision-language Dense model incorporates a linear attention mechanism, delivering fast response times while balancing inference speed and performance. Its overall capabilities are comparable to those of the Qwen3.5-122B-A10B. - [qwen3.5-35b-a3b](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/qwen3-5-35b-a3b.md): The Qwen3.5 Series 35B-A3B is a native vision-language model designed with a hybrid architecture that integrates linear attention mechanisms and a sparse mixture-of-experts model, achieving higher inference efficiency. Its overall performance is comparable to that of the Qwen3.5-27B. - [qwen3.5-397b-a17b](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/qwen3-5-397b-a17b.md): The Qwen3.5 series 397B-A17B native vision-language model is built on a hybrid architecture that integrates a linear attention mechanism with a sparse mixture-of-experts model, achieving higher inference efficiency. It delivers state-of-the-art performance comparable to leading-edge models across a wide range of tasks, including language understanding, logical reasoning, code generation, agent-based tasks, image understanding, video understanding, and graphical user interface (GUI) interactions. With its robust code-generation and agent capabilities, the model exhibits strong generalization across diverse agent. - [qwen3-235b-a22b](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/qwen3-235b-a22b.md): Achieves effective integration of thinking and non-thinking modes, allowing mode switching during conversations. Its reasoning capability significantly surpasses QwQ, and its general capability markedly exceeds Qwen2.5-72B-Instruct, reaching the SOTA level in the same scale industry. - [qwen3-235b-a22b-instruct-2507](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/qwen3-235b-a22b-instruct-2507.md): Open-source Qwen3 non-thinking model; compared to the previous version (Qwen3-235B-A22B) shows slight improvements in subjective creativity and model safety. - [qwen3-235b-a22b-thinking-2507](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/qwen3-235b-a22b-thinking-2507.md): Open-source Qwen3 thinking model; compared to the previous version (Qwen3-235B-A22B) shows major improvements in logical ability, general capabilities, knowledge enhancement, and creativity, suitable for high-difficulty, strong-thinking scenarios. - [qwen3-30b-a3b](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/qwen3-30b-a3b.md): Achieves effective integration of thinking and non-thinking modes, allowing mode switching during conversations. Its reasoning capability rivals QwQ-32B with a smaller parameter size, and its general capability significantly surpasses Qwen2.5-14B, reaching the SOTA level in the same scale industry. - [qwen3-30b-a3b-instruct-2507](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/qwen3-30b-a3b-instruct-2507.md): Open-source Qwen3 non-thinking model; compared to the previous version (Qwen3-30B-A3B) shows major improvements in Chinese, English, and overall multilingual general capabilities. Optimized for subjective open-ended tasks, delivering responses significantly more aligned with user preferences and more helpful. - [qwen3-30b-a3b-thinking-2507](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/qwen3-30b-a3b-thinking-2507.md): Open-source Qwen3 thinking model; compared to the previous version (Qwen3-30B-A3B) excels in complex thinking tasks, including logic, math, science, code, and other challenging scenarios; instruction following, text understanding, and multilingual translation capabilities significantly improved. - [qwen3-32b](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/qwen3-32b.md): Achieves effective integration of thinking and non-thinking modes, allowing mode switching during conversations. Its reasoning capability significantly surpasses QwQ, and its general capability markedly exceeds Qwen2.5-32B-Instruct, reaching the SOTA level in the same scale industry. - [qwen3-14b](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/qwen3-14b.md): Achieves effective integration of thinking and non-thinking modes, allowing mode switching during conversations. Its reasoning capability reaches the SOTA level in the same scale industry, and its general capability significantly surpasses Qwen2.5-14B. - [qwen3-8b](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/qwen3-8b.md): Achieves effective integration of thinking and non-thinking modes, allowing mode switching during conversations. Its reasoning capability reaches the SOTA level in the same scale industry, and its general capability significantly surpasses Qwen2.5-7B. - [qwen3-next-80b-a3b-instruct](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/qwen3-next-80b-a3b-instruct.md): A new generation of open-source, non-thinking mode model powered by Qwen3. This version demonstrates superior Chinese text understanding, augmented logical reasoning, and enhanced capabilities in text generation tasks over the previous iteration (Qwen3-235B-A22B-Instruct-2507). - [qwen3-next-80b-a3b-thinking](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/qwen3-next-80b-a3b-thinking.md): A new generation of Qwen3-based open-source thinking mode models. This version offers improved instruction following and streamlined summary responses over the previous iteration (Qwen3-235B-A22B-Thinking-2507). - [deepseek-v4-pro](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/deepseek-v4-pro.md): A flagship MoE large model with 1.6 trillion parameters and 49 billion activated parameters, natively supporting context lengths of up to one million tokens. Trained on a vast corpus of high-quality data, it excels in advanced mathematical reasoning, complex logical inference, specialized coding, and deep analysis of long-form text, making it well-suited for demanding applications such as cutting-edge research, sophisticated office workflows, and advanced AI agents. - [deepseek-v4-pro-us](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/deepseek-v4-pro-us.md): A flagship MoE large model with 1.6 trillion parameters and 49 billion activated parameters, natively supporting context lengths of up to one million tokens. Trained on a vast corpus of high-quality data, it excels in advanced mathematical reasoning, complex logical inference, specialized coding, and deep analysis of long-form text, making it well-suited for demanding applications such as cutting-edge research, sophisticated office workflows, and advanced AI agents. - [deepseek-v4-flash](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/deepseek-v4-flash.md): A highly efficient, lightweight MoE model with 284 billion parameters in total and 13 billion activated parameters, natively supporting context windows of up to one million tokens. It offers fast inference speed, low latency, and cost-effective invocation, delivering well-balanced overall performance. Designed for high-concurrency, lightweight workloads, it is ideally suited for common, essential use cases such as everyday dialogue, content creation, basic RAG applications, and batch text processing. - [deepseek-v4-flash-us](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/deepseek-v4-flash-us.md): A highly efficient, lightweight MoE model with 284 billion parameters in total and 13 billion activated parameters, natively supporting context windows of up to one million tokens. It offers fast inference speed, low latency, and cost-effective invocation, delivering well-balanced overall performance. Designed for high-concurrency, lightweight workloads, it is ideally suited for common, essential use cases such as everyday dialogue, content creation, basic RAG applications, and batch text processing. - [deepseek-v3.2](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/deepseek-v3-3.md): DeepSeek-V3.2 is the official release of a model that incorporates DeepSeek Sparse Attention—a sparse attention mechanism. It’s also the first model launched by DeepSeek that integrates reasoning into tool usage, supporting both reasoning-enabled and non-reasoning tool calls. - [deepseek-v3.2-exp](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/deepseek-v3-2-exp.md): Experimental version introducing DeepSeek Sparse Attention (a sparse attention mechanism), exploring optimization and validation for training and inference efficiency on long texts. - [deepseek-v3.1](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/deepseek-v3-1.md): A hybrid inference architecture model supporting both thinking mode and non-thinking mode, featuring higher reasoning efficiency and stronger agent capabilities. - [deepseek-v3](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/deepseek-v3.md): A self-developed Mixture-of-Experts (MoE) model with 671B parameters (activating 37B), pre-trained on 14.8T tokens. Demonstrates excellent capabilities in long-text processing, coding, mathematics, encyclopedic knowledge, and Chinese language tasks. - [deepseek-r1](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/deepseek-r1.md): A large language model enhanced through extensive reinforcement learning during post-training, achieving strong reasoning capabilities with minimal labeled data. Performs well in mathematics, coding, and natural language reasoning tasks. - [deepseek-r1-distill-llama-70b](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/deepseek-r1-distill-llama-70b.md): A distilled large language model based on Llama-3.1-70B, trained using DeepSeek R1's outputs. - [deepseek-r1-distill-qwen-1.5b](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/deepseek-r1-distill-qwen-1-5b.md): A distilled large language model based on Qwen2.5-Math-1.5B, trained using DeepSeek R1's outputs. - [deepseek-r1-distill-qwen-14b](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/deepseek-r1-distill-qwen-14b.md): A distilled large language model based on Qwen2.5-14B, trained using DeepSeek R1's outputs. - [deepseek-r1-distill-qwen-32b](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/deepseek-r1-distill-qwen-32b.md): A distilled large language model based on Qwen2.5-32B, trained using DeepSeek R1's outputs. - [deepseek-r1-distill-qwen-7b](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/deepseek-r1-distill-qwen-7b.md): A distilled large language model based on Qwen2.5-Math-7B, trained using DeepSeek R1's outputs. - [kimi-k3](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/kimi-k3.md): Kimi K3 is Kimi's most powerful flagship model with 2.8 trillion parameters. Built on KDA hybrid linear attention (Kimi Delta Attention) and Attention Residuals technologies, it natively supports vision understanding and features a 1 million token context window. It is the world's first open-source 3-trillion-parameter model designed for long-context programming, knowledge work, and advanced reasoning. - [kimi-k3](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/aliyun-kimi-k3.md): Kimi-K3 is Moonshot AI's most capable flagship model to date, featuring 2.8 trillion parameters. Built on the KDA hybrid linear attention mechanism (Kimi Delta Attention) and Attention Residuals, it natively supports visual understanding and offers a 1-million-token context window. As the world's first open-source model at the 3-trillion-parameter scale, it is designed for frontier intelligence scenarios such as long-horizon programming, knowledge work, and reasoning. - [kimi-k2.7-code](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/kimi-k2-7-code.md): kimi-k2.7-code is Kimi's most intelligent coding model to date. It follows instructions more reliably over long contexts and completes programming tasks with higher success rates. It supports text, image, and video inputs, along with thinking mode, conversation, and agent tasks. - [kimi-k2.6](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/kimi-k2-6.md): Kimi-k2.6 is the latest intelligent model in Kimi series, featuring enhanced long-context code generation, improved instruction following and self-correction capabilities, supporting text/image/video inputs and multiple operational modes. - [kimi-k2.5](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/kimi-k2-5.md): Kimi-k2.5 is Moonshot AI's most versatile model with native multimodal architecture, supporting visual/text inputs, thinking/non-thinking modes, and both dialogue and agent tasks. - [Moonshot-Kimi-K2-Instruct](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/moonshot-kimi-k2-instruct.md): Kimi-K2 is Moonshot's first open-source trillion-parameter MoE model in China, featuring 32B activated parameters with exceptional coding and tool-calling capabilities. - [kimi-k2-thinking](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/kimi-k2-thinking.md): Kimi-k2-thinking is a general agentic reasoning model developed by Moonshot AI, specialized in deep reasoning and multi-step tool integration to solve complex problems. - [glm-5.2-fast-preview](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/glm-5-2-fast.md): GLM-5.2-Fast-Preview is the high-speed variant of Zhipu AI's GLM-5.2, with 1M context and capabilities on par with the standard version. Inference-optimized to deliver 1.5–2× the output TPS, it fits latency-sensitive use cases such as real-time chat, multi-turn agents, and streaming code generation. - [glm-5.2](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/glm-5-2.md): GLM-5.2 is the latest flagship model from Zhipu AI, designed for long-horizon tasks with support for an ultra-long 1M context window. It features powerful logical reasoning, long-text comprehension, and code generation capabilities, balancing performance with inference efficiency. It excels across multi-task benchmarks and is well-suited for intelligent interaction, enterprise applications, and development assistance scenarios. - [glm-5.2-us](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/glm-5-2-us.md): GLM-5.2 is the latest flagship model from Zhipu AI, designed for long-horizon tasks with support for an ultra-long 1M context window. It features powerful logical reasoning, long-text comprehension, and code generation capabilities, balancing performance with inference efficiency. It excels across multi-task benchmarks and is well-suited for intelligent interaction, enterprise applications, and development assistance scenarios. - [glm-5.1](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/glm-5-1.md): GLM-5.1 is a model developed by Zhipu AI, specifically designed for long-horizon tasks. It has 744 billion parameters, supports an ultra-long context of 200k tokens, and can generate up to 128k tokens in a single response. GLM-5.1 excels in logical reasoning, long-text understanding, and code generation, while balancing performance with inference efficiency. It delivers outstanding results across multiple multi-task benchmarks and is well-suited for applications such as intelligent human-computer interaction, enterprise solutions, and developer assistance. - [glm-5](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/glm-9.md): GLM-5 targets coding and agent scenarios, achieving open-source SOTA in complex systems engineering with capabilities approaching Claude Opus, built on a 744B foundation with asynchronous reinforcement learning and sparse attention. - [glm-4.7](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/glm-4-7.md): Zhipu's latest flagship with enhanced coding and multi-step reasoning capabilities, supporting long-term task planning, tool collaboration, and immersive writing/role-playing. - [glm-4.6](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/glm-4-6.md): GLM's new flagship model with comprehensive capability improvements over 4.5, featuring a 200K context window. - [MiniMax-M2.5](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/minimax-m2-5.md): MiniMax-M2.5 is MiniMax's flagship open-source large model, trained on hundreds of thousands of real-world complex scenarios through large-scale reinforcement learning, achieving or surpassing industry SOTA in productivity scenarios like programming, tool calls, search, and office tasks. - [ZHIPU/GLM-5.3](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/glm-5-3-by-zhipu.md): Directly supplied by Z.AI , this is the latest flagship model. GLM-5.3 is the most capable open-source model Z.AI has released to date. It supports a genuinely usable 1M context and continues to lead on long-horizon tasks. - [ZHIPU/GLM-5.2](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/glm-5-2-by-zhipu.md): .GLM-5.2 1M , . - [qwen3-vl-flash](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/qwen3-vl-flash.md): The Qwen3 series of small-scale visual understanding models effectively integrates thinking and non-thinking modes, delivering superior performance compared to the open-source Qwen3-VL-30B-A3B while maintaining fast response speeds. It features a comprehensive upgrade in image/video understanding, supporting ultra-long contexts such as extended videos and documents, spatial awareness, and object recognition across various domains. Equipped with 2D/3D visual localization capabilities, it is well-suited for tackling complex real-world tasks.This model version is functionally equivalent to the snapshot model qwen3-vl-flash-2026-01-22. - [qwen3-vl-flash-us](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/qwen3-vl-flash-us.md): The Qwen3 series of small-scale visual understanding models effectively integrates thinking and non-thinking modes, delivering superior performance compared to the open-source Qwen3-VL-30B-A3B while maintaining fast response speeds. It features a comprehensive upgrade in image/video understanding, supporting ultra-long contexts such as extended videos and documents, spatial awareness, and object recognition across various domains. Equipped with 2D/3D visual localization capabilities, it is well-suited for tackling complex real-world tasks. This model version is functionally equivalent to the snapshot model qwen3-vl-flash-2026-01-22-us. - [qwen3-vl-plus](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/qwen3-vl-plus.md): The Qwen3 series VL models effectively integrates thinking and non-thinking modes, achieving world-leading performance in visual agent capabilities on public benchmark datasets such as OS World. This version features comprehensive upgrades in areas like visual coding, spatial perception, and multimodal reasoning, significantly enhancing visual perception and recognition abilities, and supporting the understanding of ultra-long videos.This model version is functionally equivalent to the snapshot model qwen3-vl-plus-2025-12-19. - [qwen-vl-max](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/qwen-vl-max.md): Qwen-VL-Max is a large-scale visual language model of the Qwen series. Compared to the Plus version, it further enhances visual reasoning capabilities and instruction-following abilities, offering higher levels of visual perception and cognition. It delivers optimal performance on more complex tasks.This model version is functionally equivalent to the snapshot model qwen-vl-max-2025-08-13. - [qwen-vl-ocr](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/qwenvl-ocr.md): Qwen-VL_OCR is an OCR model trained based on Qwen-VL. It aggregates various image-text recognition, parsing, and processing tasks through a unified model approach, offering powerful image-text recognition capabilities.This model version is functionally equivalent to the snapshot model qwen-vl-ocr-2025-11-20. - [qwen-vl-plus](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/qwen-vl-plus.md): Qwen-VL-Plus is the enhanced version of the large visual language model. It significantly improves detail recognition and text recognition capabilities, supporting images with resolutions exceeding one million pixels and any aspect ratio specifications. The model delivers exceptional performance across a wide range of visual tasks.This model version is functionally equivalent to the snapshot model qwen-vl-plus-2025-08-15. - [qwen3.5-ocr](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/qwen3-5-ocr.md): Qwen3.5-OCR is an upgraded OCR model with enhanced document parsing, text localization, and key information extraction. It significantly improves extraction performance for real-world documents (e.g., ID cards, driver's licenses). - [qvq-max](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/qvq-max.md): The Tongyi Qianwen QVQ visual reasoning model supports visual input and chain-of-thought output, demonstrating stronger capabilities in mathematics, programming, visual analysis, creation, and general tasks.This model version is functionally equivalent to the snapshot model qvq-max-2025-03-25. - [qvq-plus](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/qvq-plus.md): Qwen QVQ Visual Reasoning Model Plus version supporting visual input and chain-of-thought output with enhanced capabilities in mathematics, programming, visual analysis, creation, and general tasks. - [qwen3-vl-235b-a22b-instruct](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/qwen3-vl-235b-a22b-instruct.md): The Qwen3 series VL models has been comprehensively upgraded in areas such as visual coding and spatial perception. Its visual perception and recognition capabilities have significantly improved, supporting the understanding of ultra-long videos, and its OCR functionality has undergone a major enhancement. - [qwen3-vl-235b-a22b-thinking](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/qwen3-vl-235b-a22b-thinking.md): Qwen3 series VL models feature significantly enhanced multimodal reasoning capabilities, with a particular focus on optimizing the model for STEM and mathematical reasoning. Visual perception and recognition abilities have been comprehensively improved, and OCR capabilities have undergone a major upgrade. - [qwen3-vl-30b-a3b-instruct](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/qwen3-vl-30b-a3b-instruct.md): The Instruct version of the second-largest MoE model in the Qwen3-VL series offers rapid response speeds and supports extremely long contexts like lengthy videos and documents. It includes comprehensively upgraded image/video understanding, spatial awareness, and object recognition capabilities, as well as 2D/3D visual localization, enabling it to handle intricate real-world challenges. - [qwen3-vl-30b-a3b-thinking](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/qwen3-vl-30b-a3b-thinking.md): The Thinking version of the second-largest MoE model in the Qwen3-VL series features fast response speeds and enhanced multimodal understanding and reasoning capabilities, visual agents, and support for extremely long contexts such as lengthy videos and documents. It also boasts comprehensively upgraded image/video understanding, spatial awareness, and object recognition abilities, making it well-suited for complex real-world tasks. - [qwen3-vl-32b-instruct](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/qwen3-vl-32b-instruct.md): The largest dense model in the Qwen3-VL series, in its non-inference version, delivers overall performance second only to Qwen3-VL-235B-Instruct. It excels in document recognition and comprehension, demonstrates strong spatial awareness and object identification capabilities, and achieves state-of-the-art performance in 2D visual detection and spatial reasoning. It is well-suited for complex perception tasks across a wide range of general-purpose scenarios. - [qwen3-vl-32b-thinking](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/qwen3-vl-32b-thinking.md): The largest dense model in the Qwen3-VL series, its reasoning version boasts multimodal reasoning capabilities second only to Qwen3-VL-235B-Thinking. It excels in STEM and math problem-solving, general image and video understanding, and achieves state-of-the-art performance in multimodal agent capabilities, making it ideal for complex multimodal reasoning tasks. - [qwen3-vl-8b-instruct](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/qwen3-vl-8b-instruct.md): The Instruct version of the 8B Dense model in the Qwen3-VL series requires less GPU memory and provides comprehensively upgraded image/video understanding, support for extremely long contexts like lengthy videos and documents, spatial awareness, and object recognition capabilities, making it suitable for tackling complex real-world tasks. - [qwen3-vl-8b-thinking](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/qwen3-vl-8b-thinking.md): The Thinking version of the 8B Dense model in the Qwen3-VL series consumes less GPU memory and is capable of performing multimodal understanding and reasoning. It supports extremely long contexts such as lengthy videos and documents, 2D/3D visual localization, and features comprehensively upgraded image/video understanding, spatial awareness, and object recognition abilities. - [qwen-image-3.0-pro](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/qwen-image-3-0-pro.md): Rich content: Supports input of up to 4.5k tokens and dense information layout with images-within-images, enabling complex layouts like newspapers, storyboards, menus, and exam papers to be generated in a single pass. Authentic detail: Supports precise rendering of text as small as 10px, and vividly reproduces fine details such as micro-expressions, pores, and individual strands of hair—approaching the quality of real photography. Deep knowledge: Supports native rendering of 12 languages and 20+ fonts, realistic simulation of mainstream interfaces such as web pages, games, and live streams, fully incorporating external knowledge. Qwen-Image-3.0-Pro isn't just pursuing "good looks"—it's pursuing "usefulness", making image generation a truly deployable productivity tool. - [qwen-image-3.0](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/qwen-image-3-0.md): Accurate Prompt Understanding: Supports up to 4.5k token inputs, accurately interpreting complex text-and-image prompts and generating dense layouts in one pass. Reliable Text Rendering: Delivers crisp 10px text across 12 languages and multiple fonts, making infographics and interfaces ready to use. Efficient Batch Production: Scales everyday tasks like posters, web pages, and UI screens with better cost efficiency, keeping creative output sustainable. Qwen-Image-3.0 is built not just for visual quality, but for smooth, reliable daily creation—turning image generation into sustainable content productivity. - [qwen-image-2.0-pro](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/qwen-image-2-0-pro.md): The full-featured Qwen-Image-2.0 series models integrate image generation and image editing, offering enhanced text rendering with support for 1,000-token prompts, more refined realistic textures, detailed depiction of photorealistic scenes, and stronger semantic adherence. The full-featured version delivers the strongest text rendering and most lifelike textures in the 2.0 series.This model version is functionally equivalent to the snapshot model qwen-image-2.0-pro-2026-04-22. - [qwen-image-2.0](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/qwen-image-2-0.md): The Qwen-Image-2.0 series of accelerated models integrates image generation and image editing, offering enhanced text-rendering capabilities with support for 1,000-token prompts, more realistic textures, finely detailed photorealistic scenes, and improved semantic consistency. The accelerated version effectively strikes an optimal balance between model performance and quality.This model version is functionally equivalent to the snapshot model qwen-image-2.0-2026-03-03. - [qwen-image-edit-max](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/qwen-image-edit-max.md): The Max series Qwen’s image editing models delivers more stable and versatile editing capabilities: enhanced industrial design and geometric reasoning, improved character consistency, reduced offset issues, and integrated LoRA capabilities for a wider range of image editing functions.This model version is functionally equivalent to the snapshot model qwen-image-edit-max-2026-01-16. - [qwen-image-edit-plus](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/qwen-image-edit-plus.md): The qwen series of image editing Plus models further optimizes inference performance and system stability based on the initial Edit model, significantly reducing the response time for image generation and editing. It also supports returning multiple images in a single request, greatly enhancing user experience. - [qwen-image-edit](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/model-qwen-image-edit.md): The first Qwen image editing model extends Qwen-Image's text rendering to editing tasks. It offers precise bilingual (Chinese/English) text editing, dual visual and semantic editing, and strong cross-benchmark performance. - [qwen-image-max](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/qwen-image-max.md): The Max series of qwen’s image generation model excels across a wide range of generation tasks. Compared with the Plus series, it significantly reduces the “AI-like” feel in generated images, enhancing their realism. It delivers more lifelike material textures for human subjects, finer and more detailed natural textures, and more visually appealing text rendering.This model version is functionally equivalent to the snapshot model qwen-image-max-2025-12-30. - [qwen-image-plus](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/qwen-image-plus.md): The first image generation foundation model in the Qwen series that achieves significant advances in complex text rendering and precise image editing. Experiments show strong general capabilities in both image generation and editing, with exceptional performance in text rendering, especially for Chinese. - [qwen-image](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/qwen-image.md): The first image generation foundation model in the Qwen series that achieves significant advances in complex text rendering and precise image editing. Experiments show strong general capabilities in both image generation and editing, with exceptional performance in text rendering, especially for Chinese. - [qwen-mt-image](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/qwen-mt-image.md): Qwen-MT-Image specializes in image translation, converting images across 11 languages (Chinese, English, Japanese, etc.) to target languages. It accurately preserves layout and content, supporting custom features like terminology definitions, sensitive word filtering, and product detection for flexible, accurate image localization. - [wan2.7-image-pro](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/wan2-7-image-pro.md): Wan2.7–image-pro, supports text to image, text/image to sequential images, image editing, multi-image reference generation, and interactive editing. Delivers enhanced performance in text rendering, subject consistency, and complex instruction following. - [wan2.7-image](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/wan2-7-image.md): Wan2.7 – image generation and editing, supports text to image, text/image to sequential images, image editing, multi-image reference generation, and interactive editing. Delivers enhanced performance in text rendering, subject consistency, and complex instruction following. - [wan2.6-image](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/wan2-6-image.md): Wan2.6 Image, An all-round image generation model that supports joint text–image reasoning, multi-image creative fusion, commercial-grade consistency, aesthetic style transfer, and precise control of framing and lighting, significantly enhancing consistency, controllability, and expressiveness in image generation. - [wan2.6-t2i](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/wan2-6-t2i.md): Wan2.6 text to image, Upgraded visual quality, aesthetics, and instruction-following deliver precise style control, realistic portraits, long-text understanding, and broad historical/cultural IP coverage, enabling high-quality, highly expressive visual generation. - [wan2.5-i2i-preview](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/wan2-5-i2i.md): The upgraded Wan2.5 Preview image edit model, newly upgraded model architecture supports rich image editing capabilities via instruction control, with enhanced instruction adherence. It also enables multi-image reference generation with high consistency and demonstrates excellent text generation performance. - [wan2.5-t2i-preview](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/wan2-5-t2i.md): The upgraded Wan2.5 Preview text to image model, newly upgraded model architecture significantly enhances visual aesthetics, design sensibility, and realistic texture. It excels in precise instruction adherence, generates text proficiently in English, Chinese, and less common languages, and supports the generation of complex structured long texts, charts, and architectural diagrams. - [wan2.2-t2i-flash](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/wan2-2-t2i-flash.md): The upgraded Wan 2.2 Flash text to image model, delivers faster speed with enhanced creativity, stability, and realism. It also features stronger prompt following and native support for multiple styles. Up to 2 million pixel generation and prompt enhancement are supported as well. - [wan2.2-t2i-plus](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/wan2-2-t2i-plus.md): The upgraded Wan 2.2 Plus text to image model, delivers richer image detail with enhanced creativity, stability, and realism. It also features stronger prompt following and native support for multiple styles. Up to 2 million pixel generation and prompt enhancement are supported as well. - [wanx2.1-imageedit](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/wanx2-1-imageedit.md): Image Editing model supporting preset and command-based tasks, including global/local editing (style transfer, inpainting, expansion, super-resolution) and reference-based generation. - [wan2.1-t2i-plus](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/wan2-1-t2i-plus.md): Wan2.1 Text-to-Image Plus version, Generate more image details. Upgraded in image beauty, realism, and artistry. Stronger semantic understanding ability, rich style generalization ability, supports up to 2 million pixel generation, supports smart prompt rewriting. - [wanx2.1-t2i-plus](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/wanx2-1-t2i-plus.md): Wanxiang 2.1 - Text-to-Image - Plus. Upgraded image generation with enhanced aesthetics, realism, and artistry. Improved semantic understanding, style generalization, and support for up to 2MP resolution. Features smart prompt rewriting and richer visual details. - [wan2.1-t2i-turbo](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/wan2-1-t2i-turbo.md): Wan2.1 Text-to-Image Turbo version, faster generation speed. Upgraded in image beauty, realism, and artistry. Stronger semantic understanding ability, rich style generalization ability, supports up to 2 million pixel generation, supports smart prompt rewriting. - [wanx2.1-t2i-turbo](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/wanx2-1-t2i-turbo.md): Wanxiang 2.1 - Text-to-Image - Turbo. Accelerated generation speed with enhanced aesthetics, realism, and artistry. Maintains strong semantic understanding, style diversity, and 2MP resolution support. Includes smart prompt rewriting functionality. - [wanx2.0-t2i-turbo](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/wanx2-0-t2i-turbo.md): Enhanced Text-to-Image model specializing in realistic portraits and creative designs, upgraded in aesthetics, realism, and artistic quality with up to 2MP resolution and smart prompt rewriting support. - [z-image-turbo](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/z-image-turbo.md): Z-Image-Turbo is a highly efficient image-generation model that has topped the Artificial Analysis benchmark as the world’s No. 1 open-source text-to-image model. With just 6 billion parameters and an 8-step inference process, it generates photo-realistic images comparable to those produced by large-scale commercial models, while excelling in bilingual Chinese–English text rendering, complex semantic understanding, and diverse thematic generation. - [aitryon-parsing-v1](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/aitryon-parsing-v1.md): An image segmentation model serving as a supporting module for AI try-on OutfitAnyone, capable of segmenting model images and garment images for pre/post-processing of try-on images. - [aitryon-plus](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/aitryon-plus.md): A high-quality virtual try-on image generation model that produces try-on effects with enhanced image clarity, garment texture details, and logo reconstruction compared to aitryon. Requires longer generation time, suitable for non-time-critical scenarios. - [image-erase-completion](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/image-erase-completion.md): Image erasure and completion tool removing specified elements (people, objects, text, watermarks) while preserving backgrounds using computer vision and AIGC inpainting techniques. - [image-instance-segmentation](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/model-image-instance-segmentation.md): Person instance segmentation employs detection and segmentation techniques to identify objects in images and generate pixel-level masks for precise object boundary delineation. - [FAQ](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/faq-about-alibaba-cloud-model-studio.md): This document answers frequently asked questions about Alibaba Cloud Model Studio. - [Related agreements](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/related-agreements.md): These agreements govern your use of Model Studio. Review them before using the service. ## Changelog - [Model decommissioning policy](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/model-depreciation.md): To optimize resources and provide users with the latest, most advanced models, Alibaba Cloud Model Studio periodically retires legacy models. This topic describes the model retirement process. - [Model lifecycle and updates](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/newly-released-models.md): The following tables list model releases . For model deprecation rules and lists, see Model decommissioning policy . - [Model Platform Feature Update Announcements](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/model-release-notes.md) ## Get Started - [Get Started](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/start-using.md) - [Build a private knowledge Q&A application with zero code](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/build-knowledge-base-qa-assistant-without-coding.md): Large language models (LLMs) cannot directly answer questions about private knowledge. Model Studio lets you connect your own documents to an agent application—no code required—so users get accurate, grounded answers instead of generic or fabricated responses. ## Application Development - [Application Development](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/llm-application.md) - [Application types](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/application-introduction.md): Large language models have limitations in processing private knowledge, accessing real-time information, following fixed procedures, and planning complex tasks. To overcome these challenges, Model Studio provides two application types: agent applications, workflow applications. By integrating capabilities such as knowledge base retrieval, external tool calls, and memory, Model Studio lets you build AI applications that solve business problems. - [Agent applications](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/single-agent-application.md): Alibaba Cloud Model Studio agent applications let you connect an LLM to external tools and knowledge bases without code, extending the model's capabilities beyond its built-in limits. - [Workflow application](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/workflow-application.md): A workflow breaks down complex tasks into ordered steps to reduce system complexity. In Alibaba Cloud Model Studio, workflows let you combine nodes, such as large models, APIs, and Function Compute, to reduce coding costs. ## Prompt - [Prompt template overview](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/prompt-template.md): Prompt templates let you separate the fixed structure of a prompt from its dynamic variables, creating reusable templates for unified management, optimization, and efficient prompt generation. - [Custom prompt templates](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/prompt-custom-template.md): Building high-quality prompts from scratch is time-consuming and can lead to inconsistent results. Custom prompt templates in Model Studio let you create structured, reusable templates. Use them to design, manage, and optimize prompts for stable, high-quality output. - [Optimize model outputs with a prompt sample library](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/prompt-sample-optimization.md): General-purpose models can struggle with precise, formatted answers for specialized tasks. Using few-shot learning, retrieve relevant examples from predefined high-quality Q&A pairs to guide the model toward more accurate and consistent responses. Use this feature when responses must follow a strict style: customer service, domain-specific Q&A, and formatted content generation. - [Automatic prompt optimization](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/optimize-prompt.md): Crafting high-quality prompts can be time-consuming and require specialized knowledge. Automatic prompt optimization in Alibaba Cloud Model Studio simplifies this process. It uses a model to analyze and rewrite your original prompt, generating a version with better structure, clearer instructions, and more reliable results. This enhances the performance and consistency of your model-based applications. ## Application Data - [Application Data](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/application-data.md) - [Data import](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/data-import-instructions.md): To build a knowledge base, import your source data. ## Knowledge Base (RAG) - [Knowledge base](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/rag-knowledge-base.md): A knowledge base supplements an LLM with private data and up-to-date information. Using retrieval-augmented generation (RAG), the LLM improves answer accuracy by first retrieving relevant content from the knowledge base before generating a response. - [RAG performance optimization](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/rag-optimization.md): If you encounter incomplete knowledge retrieval or inaccurate content with the retrieval-augmented generation (RAG) feature in Alibaba Cloud Model Studio, refer to the suggestions and examples in this topic to improve RAG performance. - [Knowledge base API](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/rag-knowledge-base-api-guide.md): The Alibaba Cloud Model Studio knowledge base provides open APIs that enable you to integrate with your existing business systems, automate operations, and address complex retrieval needs. - [Knowledge base quotas and limits](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/rag-knowledge-base-specifications.md) - [knowledge retrieval](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/rag-knowledge-retrieval.md): The Knowledge Retrieval service supports searching a single knowledge base or performing a multi-knowledge-base federated search to help you find relevant content from your private enterprise knowledge bases. You can create and configure retrieval services in the console or integrate them into your applications using the API. ## Plugin - [Plugin](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/plug-in.md) - [Plugin overview](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/plug-in-overview.md): Models have limitations: they cannot access the latest information, are prone to hallucination, and struggle with precise calculations. Integrate plugins to address these limitations. - [Official plugins](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/plugins.md): While large language models have powerful natural language processing capabilities, they may need additional features for specific tasks, such as performing web searches or processing images. Model Studio provides a range of official plugins. You can select plugins to enhance the model's capabilities and expand its use cases. - [Custom plugins](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/custom-plug-ins.md): This document guides you through creating, debugging, and using custom plugins to integrate the APIs you need. ## Application Calling - [Application Calling](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/bailian-application-calling.md) - [Call applications](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/application-calling-guide.md): You can integrate Model Studio applications—such as agents, workflows, or agent orchestrations—into your business systems using the DashScope SDK or HTTP. - [Custom application parameter pass-through](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/pass-through-of-application-parameters.md): Pass custom parameters, primarily to custom plugins and custom nodes, when calling Alibaba Cloud Model Studio's Agent Application and Workflow Application (which replace agent orchestration applications). ## Permissions - [Permission management](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/application-permission-management-overview.md): Alibaba Cloud Model Studio offers granular access control at the console and model level, helping organizations manage users across multiple regions. ## Assistant API (Deprecated) - [Getting started with Assistant API (Deprecated)](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/quick-start-of-assistant-api.md): The Assistant API provides a set of development tools to help you easily manage conversation messages and call tools. This topic uses the example of building a painting assistant from scratch to help you quickly learn the basic encoding methods of the Assistant API. - [Tool calling overview (Deprecated)](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/tool-calling-overview.md) - [Function calling (Deprecated)](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/function-calling.md): The Assistant API supports function calling. This feature allows an agent to automatically call external functions to perform tasks, such as translating text. This topic uses a simple "Translation Agent" example to help you quickly understand the basics of function calling. ## Support - [FAQ](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/application-faq.md): Common questions about Model Studio applications and data management. ## API Usage - [Obtain an API key](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/get-api-key.md): Before you use models or applications in Alibaba Cloud Model Studio, you must obtain an API key for authentication. - [Install the SDK](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/install-sdk.md): Model Studio provides official DashScope SDKs (Python, Java) and supports OpenAI SDKs (Python, Node.js, Java, Go) for OpenAI-compatible API calls. - [Use Model Studio CLI](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/use-model-studio-cli.md): Alibaba Cloud Model Studio CLI is a command-line tool built by Alibaba Cloud Model Studio specifically for AI Agents. With a single installation command and authentication setup, you can integrate the AI capabilities of Model Studio into various AI tools. - [Error codes](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/error-code.md): This topic describes common error messages you might encounter when using Alibaba Cloud Model Studio and provides solutions. ## Text Generation - [Text Generation API Reference](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/qwen-api-reference.md) - [OpenAI compatible - Chat](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/qwen-api-via-openai-chat-completions.md): You can call models using the OpenAI compatible Chat API. This document describes the input and output parameters and provides call examples. - [Create a response](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/qwen-api-via-openai-responses.md): Use the OpenAI-compatible Responses API to call the Qwen model. This topic describes the input and output parameters and provides a call example. - [Retrieve a response](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/retrieve-a-response.md): Retrieve a completed model response by its Response ID. - [Delete a response](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/delete-a-response.md): Deletes a stored model response based on its response ID. - [List input items](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/list-input-items.md): Returns the input items used to generate the specified Response. For multi-turn conversations chained by previous_response_id , the returned list also includes user inputs and assistant replies from earlier turns. A Response ID can be queried only when the create request was sent with store=true . - [Anthropic-compatible Messages](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/anthropic-api-messages.md): Migrate your Anthropic application to Model Studio by changing three settings. This topic covers the request and response parameters with code examples. - [DashScope API Reference](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/qwen-api-via-dashscope.md): You can call Qwen models using the DashScope API. This topic describes the input and output parameters and provides call examples. ## Image Generation - [Qwen Image Generation and Editing 3.0 API Reference](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/qwen-image-generation-and-editing-api-reference.md): The Qwen Image Generation and Editing 3.0 model supports both text-to-image (T2I) and image-to-image/image editing (I2I). It can generate images directly from text prompts or edit images based on 1-3 reference images combined with editing instructions. - [Qwen-Image API reference](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/qwen-image-api.md): Qwen-Image is a general-purpose image generation model that supports multiple artistic styles and excels at complex text rendering . It handles multi-line layouts, paragraph-level text generation, and fine-grained detail rendering. - [Qwen-Image Edit API reference](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/qwen-image-edit-api.md): Qwen-Image Edit supports multi-image input and output. Edit text within images, add, remove, or move objects, change subject poses, transfer styles, and enhance details — all through natural language prompts. - [Qwen-MT-Image API reference](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/qwen-mt-image-api.md): Qwen-MT-Image accurately translates text in images while preserving the original layout. The model also supports domain hints, sensitive word filtering, and terminology intervention. - [Wan text-to-image V2 API reference](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/text-to-image-v2-api-reference.md): The Wan text-to-image model generates images from text prompts, supporting artistic styles and realistic photographic effects. - [Wan2.7 - image generation and editing](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/wan-image-generation-and-editing-api-reference.md): Wan2.7-Image supports text-to-image, text-to-image-set, image-to-image-set, image editing, and multi-image reference generation. - [Wan2.6 - image generation and editing](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/wan-image-generation-api-reference.md): Wan2.6 supports multi-image input, image editing, and interleaved text-image output. - [Wanxiang – General Image Editing 2.5](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/wan2-5-image-edit-api-reference.md): The Wanxiang General Image Editing wan2.5 model edits and fuses images from text instructions alone, maintaining subject consistency across edits. - [Wan2.1 - general image editing API reference](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/wanx-image-edit-api-reference.md): This topic describes the input and output parameters for the Wan - general image editing model. - [Z-Image API reference](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/z-image-api-reference.md): A lightweight text-to-image model for fast generation, with Chinese and English text rendering, and flexible resolutions. - [Image erase completion API reference](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/image-erase-completion-api-reference.md): This document details the parameters for the image erase completion model. This model removes one or more elements from an image, such as people, pets, objects, text, or watermarks, while preserving the background. You can specify the areas to remove using a mask image. - [OutfitAnyone](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/outfitanyone.md): OutfitAnyone includes try-on models and auxiliary models for virtual try-on scenarios, from quick image generation and detail refinement to partial replacement. - [OutfitAnyone-Plus API reference](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/aitryon-plus-api.md): Aitryon-plus delivers higher image clarity, better fabric texture, and more accurate logos, but takes longer to process. - [OutfitAnyone-Parsing API reference](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/aitryon-parsing-api.md): OutfitAnyone-Parsing segments clothing areas such as tops, bottoms, dresses, or jumpsuits from model images or OutfitAnyone-generated images. Use it with OutfitAnyone for partial try-on or to get clothing coordinates . - [FAQ](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/image-faq.md): Common questions about image APIs in Alibaba Cloud Model Studio, covering debugging, billing, rate limits, and API errors. ## Video Generation - [HappyHorse text-to-video API reference](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/happyhorse-text-to-video-api-reference.md): Generate physically realistic, motion-smooth video from text prompts with the HappyHorse model. - [HappyHorse image-to-video (first frame) API reference](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/happyhorse-image-to-video-api-reference.md): Generate realistic, smooth-motion videos from a first-frame image and an optional text prompt using the HappyHorse model. - [HappyHorse - reference-to-video API reference](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/happyhorse-reference-to-video-api-reference.md): The HappyHorse reference-to-video model lets you provide multiple reference images and a text prompt to generate a video that combines subjects from the images into a scene based on the prompt. - [HappyHorse video editing API reference](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/happyhorse-video-edit-api-reference.md): The HappyHorse video editing model takes a video and a reference image as input, and performs editing tasks such as style transfer and local replacement based on text instructions. - [Wan3.0 - Video Generation API Reference](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/wan3-video-generation-api-reference.md): Wan3.0 is an All-in-One reference-based video generation model that supports Text-to-Video , Image-to-Video (first frame/first-last frame), and Reference-based Video Generation . It can generate videos up to 30 seconds long at 30fps. Currently in preview . - [Wan 2.7 - image-to-video API](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/image-to-video-general-api-reference.md): The Wan 2.7 image-to-video model supports multimodal input (text, images, audio, and video) and performs three tasks: first-frame-to-video, first-and-last-frame-to-video, and video continuation . - [Wan2.7 - text-to-video API reference](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/text-to-video-api-reference.md): The Wan text-to-video model generates smooth videos from text prompts . - [Wan - reference-to-video API reference](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/wan-video-to-video-api-reference.md): Wan-R2V accepts multimodal input (images, videos, and audio) to generate videos featuring one or more characters while preserving their appearance and voice across scenes. - [Wan2.7 - video editing API reference](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/wan-video-editing-api-reference.md): The Wan 2.7 video editing model accepts multimodal inputs (text, images, and videos) for instruction-based editing and style transfer . - [Wan - image-to-video - first frame API reference (2.1-2.6)](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/legacy-image-to-video-api-reference.md): The Wan image-to-video model generates a smooth video from a first-frame image and a text prompt . - [Wanxiang - Image-to-Video - Video Effects](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/wanx-video-effects.md): A video effect template is a preset for video generation effects. Provide a first-frame image and a template value to generate a video with a specific dynamic effect -- such as "Magic Levitation" or "Squeeze". This feature is suitable for social media, marketing promotions, and interactive entertainment. - [Wan - text-to-video API reference (2.1-2.6)](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/legacy-wan-text-to-video-api-reference.md): The Wan text-to-video model generates smooth videos from text prompts . - [Wan - reference-to-video (2.6)](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/legacy-wan-reference-to-video-api-reference.md): The Wan reference-to-video model accepts multimodal input and generates single-character or multi-character interaction videos using people or objects as protagonists. - [Wan -image-to-video-first and last frames API reference(2.2)](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/legacy-image-to-video-by-first-and-last-frame-api-reference.md): The Wan 2.2 model generates a smoothly transitioning video from a first frame , a last frame, and a text prompt . - [Wan - video editing (2.1)](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/legacy-wanx-vace-api-reference.md): The Wan 2.1 unified video editing model supports multiple input modalities, including text, images, and videos, for a wide range of video generation and editing tasks. - [Wan image-to-action API reference](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/wan-animate-move-api.md): Animate a character image by transferring actions from a reference video. - [Wan - video character swap API reference](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/wan-animate-mix-api.md): Replaces the main character in a video with a character from an image while preserving the original scene, lighting, and tone for seamless integration. - [Wanxiang - Digital Human](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/wan-s2v-overview.md): The wan2.2-s2v digital human model generates natural-looking speaking, singing, or performing videos from a single image and an audio file . It supports any aspect ratio and works with portrait, full-body, or half-body images. - [wan2.2-s2v image detection API reference](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/wan-s2v-detect-api.md): wan2.2-s2v-detect checks whether an input image meets wan2.2-s2v input specifications. - [Digital human wan2.2-s2v video generation API reference](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/wan-s2v-api.md): The digital human wan2.2-s2v model generates natural talking, singing, or performing videos based on a single image and an audio file . - [AnimateAnyone: Generate dance videos from images](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/animateanyone-quick-start.md): AnimateAnyone generates videos of a character in motion from a character image and a motion template. It includes three independent models: "AnimateAnyone-detect", "AnimateAnyone-template", and "AnimateAnyone". These models provide capabilities for character image compliance detection, motion template generation, and character video generation. - [AnimateAnyone image detection API reference](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/animate-anyone-detect-api.md) - [AnimateAnyone action template generation API reference](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/animate-anyone-template-api.md): This model extracts character movements from videos to generate action templates for AnimateAnyone. This topic describes the template generation API. - [AnimateAnyone video generation API reference](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/animateanyone-video-generation-api.md): The video generation model of AnimateAnyone generates videos from an action template and an image. This topic describes how to call the video generation API. - [Image-to-Singing-and-Acting Video – EMO](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/emo-quick-start.md): EMO generates dynamic portrait videos from a portrait image and human speech audio file. It consists of two models: EMO-detect verifies input image requirements, and EMO generates the video. - [EMO image detection API reference](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/emo-detect-api.md): The emo-detect-v1 model validates whether a portrait image meets the requirements for EMO video generation . Call this API to check your image before generating a video, and retrieve the bounding box coordinates required by the EMO video generation API. - [EMO video generation API reference](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/emo-api.md): Generate animated face videos from portrait images and voice audio. Submit an image and audio file to receive a video with the face animated to match the speech. - [Image to broadcast video - LivePortrait](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/liveportrait-quick-start.md): LivePortrait generates dynamic portrait videos from portrait images and audio files with a human voice. It includes two models: `LivePortrait-detect` (validates image compliance) and `LivePortrait` (generates videos). - [LivePortrait image detection API reference](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/liveportrait-detect-api.md): The LivePortrait-detect model checks whether input images meet LivePortrait specifications. This document describes how to call the detection API. - [LivePortrait video generation API reference](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/liveportrait-api.md): Generate dynamic portrait videos from portrait images processed by LivePortrait-detect and human voice audio files. This document describes the video generation API. - [Lip-sync replacement for videos - VideoRetalk](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/videoretalk.md): VideoRetalk synchronizes lip movements with audio, generating a new video from an input video and an audio file. - [VideoRetalk API reference](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/videoretalk-api.md): Use the VideoRetalk API to generate videos with synchronized lip movements by replacing a person's speech with a provided audio track. - [Image to emoji video - Emoji](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/emoji-quick-start.md): Emoji creates expressive videos from a portrait or half-body image and a preset dynamic template. - [Emoji image detection API reference](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/emoji-detect-api.md): Detects whether an input image meets Emoji model requirements. If detection passes, the model returns face area coordinates (bbox_face) and extended dynamic area coordinates (ext_bbox_face) for video generation. - [Emoji video generation API reference](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/emoji-api.md): The emoji-v1 model generates facial emoji videos from portrait images and preset template IDs . - [Video style transform API Reference](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/video-style-transform-api-reference.md): Transform input videos into artistic styles while preserving smooth motion and content coherence. Supports eight styles: Japanese manga, American comic, fresh comic, 3D cartoon, Chinese cartoon, paper art, simple illustration, and Chinese ink wash painting. ## Audio - [Qwen-Audio-3.0-ASR-Flash-Streaming/Fun-ASR-Realtime real-time speech recognition WebSocket API reference](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/fun-asr-realtime-websocket-api.md): This topic describes the service endpoint, request headers, and interaction flow for accessing the Qwen-Audio-3.0-ASR-Flash-Streaming/Fun-ASR-Realtime real-time speech recognition service over a WebSocket connection. - [Qwen-Audio-3.0-ASR-Flash-Streaming/Fun-ASR-Realtime real-time speech recognition client events](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/fun-asr-client-events.md): This topic describes the client events that the client sends to the server over WebSocket in the Qwen-Audio-3.0-ASR-Flash-Streaming/Fun-ASR-Realtime real-time speech recognition service, including the data structures and field definitions for run-task (start a task), and finish-task (end a task). - [Qwen-Audio-3.0-ASR-Flash-Streaming/Fun-ASR-Realtime real-time speech recognition server-side events](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/fun-asr-server-events.md): The Qwen-Audio-3.0-ASR-Flash-Streaming/Fun-ASR-Realtime real-time speech recognition service pushes server-side events to the client over WebSocket. This topic describes the data structures and field descriptions of the four event types: task-started, result-generated, task-finished, and task-failed. - [Qwen-Audio-3.0-ASR-Flash-Streaming/Fun-ASR-Realtime Python SDK](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/fun-asr-realtime-python-sdk.md): This topic describes the parameters and interfaces of the Python SDK for the Qwen-Audio-3.0-ASR-Flash-Streaming/Fun-ASR-Realtime real-time speech recognition model. - [Qwen-Audio-3.0-ASR-Flash-Streaming/Fun-ASR-Realtime Java SDK provides interfaces for synchronous and streaming speech recognition](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/fun-asr-realtime-java-sdk.md): This topic describes the parameters and interfaces of the Java SDK for Qwen-Audio-3.0-ASR-Flash-Streaming/Fun-ASR-Realtime real-time speech recognition. - [Interaction flow for real-time speech recognition (Qwen-ASR-Realtime)](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/qwen-asr-realtime-interaction-process.md): Qwen-ASR-Realtime receives audio streams and transcribes speech in real time over WebSocket. The service supports two interaction modes: VAD mode and Manual mode . - [Client events for Qwen-ASR-Realtime](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/qwen-asr-realtime-client-events.md): This page documents client-to-server events for the Qwen-ASR Realtime WebSocket API. Each section covers an event type, its parameters, and server responses. - [Server events for Qwen-ASR-Realtime](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/qwen-asr-realtime-server-events.md): This topic describes the events that the server sends to the client during a WebSocket session with the Qwen-ASR-Realtime API. - [Qwen-ASR-Realtime Python SDK - API reference](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/qwen-asr-realtime-python-sdk.md): Stream audio to Qwen-ASR-Realtime over WebSocket and receive real-time transcription results via the DashScope Python SDK. - [Qwen-ASR-Realtime Java SDK - API reference](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/qwen-asr-realtime-java-sdk.md): Use the DashScope Java SDK to call Qwen-ASR-Realtime. - [WebSocket API](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/websocket-for-paraformer-real-time-service.md): Access the Paraformer real-time speech recognition service over a WebSocket connection. This topic describes the service endpoints, request headers, and interaction flow. - [Paraformer real-time speech recognition client events](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/paraformer-client-events.md): Two WebSocket client events control a Paraformer real-time speech recognition task: run-task starts the task with the model and audio settings, and finish-task ends the task after the audio stream completes. This page describes the message structure and field semantics of both events. - [Server-sent events for real-time speech recognition (Paraformer)](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/paraformer-server-events.md): Reference for the server-sent events that the Paraformer real-time speech recognition service pushes to clients over WebSocket. This topic documents the data structure and field semantics of the four event types: task-started, result-generated, task-finished, and task-failed. - [Python SDK](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/paraformer-real-time-speech-recognition-python-sdk.md): The parameters and interfaces of the Paraformer real-time speech recognition Python SDK. - [Paraformer Real-time Speech Recognition Java SDK](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/paraformer-real-time-speech-recognition-java-sdk.md): This topic describes the parameters and interface details of the Paraformer real-time speech recognition Java SDK. - [Qwen-Audio-3.0-ASR-Flash-Filetrans/Fun-ASR HTTP API for non-real-time speech recognition](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/fun-asr-recorded-speech-recognition-http-api.md): This topic describes the parameters and interface details of the HTTP API for non-real-time speech recognition with Qwen-Audio-3.0-ASR-Flash-Filetrans and Fun-ASR. - [Qwen-Audio-3.0-ASR-Flash-Filetrans/Fun-ASR non-real-time speech recognition Python SDK](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/funauidio-asr-recorded-speech-recognition-python-sdk.md): This topic describes the parameters and interfaces of the Qwen-Audio-3.0-ASR-Flash-Filetrans/Fun-ASR non-real-time speech recognition Python SDK. - [Qwen-Audio-3.0-ASR-Flash-Filetrans/Fun-ASR non-real-time speech recognition Java SDK](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/fun-asr-recorded-speech-recognition-java-sdk.md): This topic describes the parameters and API details of the Qwen-Audio-3.0-ASR-Flash-Filetrans/Fun-ASR non-real-time speech recognition Java SDK. - [Non-real-time speech recognition (Qwen-Audio-3.0-ASR-Flash/Fun-ASR-Flash) API reference](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/non-real-time-speech-recognition-for-fun-asr-flash.md): This topic describes the parameters and interface details of the Qwen-Audio-3.0-ASR-Flash/Fun-ASR-Flash non-real-time speech recognition HTTP API. - [Non-real-time speech recognition (Qwen-ASR) API reference](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/qwen-asr-api-reference.md): Input and output parameters for the Qwen-ASR model. Call the API using the OpenAI compatible or DashScope protocol. - [Paraformer non-real-time speech recognition HTTP API](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/paraformer-recorded-speech-recognition-restful-api.md): The parameters and API details for Paraformer non-real-time speech recognition HTTP API. - [Paraformer non-real-time speech recognition Python SDK](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/paraformer-recorded-speech-recognition-python-sdk.md): Use the Paraformer Python SDK to transcribe audio and video files with the DashScope API. - [Paraformer non-real-time speech recognition Java SDK](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/paraformer-recorded-speech-recognition-java-sdk.md): This topic describes the parameters and interface details of the Paraformer non-real-time speech recognition Java SDK. - [Best practices](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/paraformer-best-practices.md) - [Custom vocabulary HTTP API reference](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/vocabulary-http-api.md): Manage custom vocabularies through HTTP APIs, including creating, listing, getting, updating, and deleting vocabularies. - [Custom hotwords Python SDK reference](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/vocabulary-python-sdk.md): Use the Python SDK to create, list, query, update, and delete custom vocabularies for speech recognition. - [Custom hotwords Java SDK reference](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/vocabulary-java-sdk.md): Use the Java SDK to create, query, update, and delete custom vocabularies for speech recognition. - [Qwen-Audio-TTS/CosyVoice speech synthesis WebSocket API](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/cosyvoice-websocket-api.md) - [Qwen-Audio-TTS/CosyVoice client events](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/cosyvoice-client-events.md) - [Qwen-Audio-TTS/CosyVoice server-side events](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/cosyvoice-server-events.md) - [Qwen-Audio-TTS/CosyVoice speech synthesis Java SDK](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/cosyvoice-java-sdk.md): Synthesize speech with Qwen-Audio-TTS/CosyVoice using the DashScope Java SDK. - [Qwen-Audio-TTS/CosyVoice speech synthesis Python SDK](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/cosyvoice-python-sdk.md): Use the DashScope Python SDK to integrate Qwen-Audio-TTS/CosyVoice real-time speech synthesis into your application through non-streaming, one-way streaming, or bidirectional streaming modes. - [WebSocket API for Qwen-TTS real-time synthesis](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/interactive-process-of-qwen-tts-realtime-synthesis.md): Connect to the Qwen-TTS real-time speech synthesis service over WebSocket. Covers the service endpoint, request headers, and interaction flow. - [Client events](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/qwen-tts-realtime-client-events.md): Client events are JSON messages sent over a WebSocket connection to control the Qwen-TTS Realtime API session -- configure voice settings, stream text for synthesis, and signal completion. - [Server events](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/qwen-tts-realtime-server-events.md): Server events for the Qwen-TTS-Realtime API. - [Python SDK](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/qwen-tts-realtime-python-sdk.md) - [Java SDK](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/qwen-tts-realtime-java-sdk.md): Key interfaces and request parameters for Qwen real-time speech synthesis DashScope Java SDK. - [Qwen-TTS non-real-time speech synthesis API reference](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/qwen-tts-api.md): Request parameters and response fields for the non-realtime speech synthesis (Qwen-TTS) API. - [Voice cloning HTTP API reference](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/voice-clone-design-http-api.md): Use the HTTP API to create, list, query, update, and delete cloned voices. - [Voice cloning Java SDK reference](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/voice-clone-java-sdk.md): Use the DashScope Java SDK to clone and manage Qwen-Audio-TTS/CosyVoice voices. - [Voice cloning Python SDK reference](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/voice-clone-python-sdk.md): Qwen-Audio-TTS/CosyVoice voice cloning is accessible through the DashScope Python SDK. - [Voice Design API reference](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/voice-design-api-references.md): Use the Voice Design HTTP API to create, list, query, and delete custom voices. - [Music generation API reference(Fun-Music)](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/fun-music-api.md): API parameters for the Fun-Music music generation model. - [Client events](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/live-translator-client-events.md): This topic describes the client events for the qwen3.5-livetranslate-flash-realtime API. - [Server events](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/live-translator-server-events.md): Server-side events for the qwen3.5-livetranslate-flash-realtime API. - [Qwen-LiveTranslate Python SDK - API reference](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/qwen-livetranslate-python-sdk.md): Use the DashScope SDK for Python to call Qwen-LiveTranslate for real-time speech translation. - [Qwen-LiveTranslate Java SDK - API reference](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/qwen-livetranslate-java-sdk.md): Translate speech in real time using the DashScope Java SDK and the qwen3.5-livetranslate-flash-realtime model. The SDK connects over WebSocket, streams audio input, and returns translated text and synthesized speech. - [Audio and video translation - Qwen API reference](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/qwen3-livetranslate-flash-api.md): qwen3-livetranslate-flash translates audio and video through the OpenAI-compatible chat completions endpoint. All requests are streamed. - [Qwen-Audio Realtime WebSocket API reference](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/fun-audiochat-realtime-websocket-api.md): The Qwen-Audio Realtime API provides real-time voice conversation capabilities over the WebSocket protocol. Clients interact with the server by sending and receiving JSON events. The API supports audio input, text input, voice activity detection (VAD), and streaming audio and text output. - [Qwen-Audio client events](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/fun-audiochat-client-events.md): Client event reference for the Qwen-Audio Realtime API. - [Server events for Qwen-Audio Realtime API](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/qwen-audio-realtime-server-events.md): Server event reference for the Qwen-Audio Realtime API. All server events include the event_id (auto-generated by the server) and type (event type) fields. ## Real-time Multimodal - [Client events](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/client-events.md): Client events for the Qwen-Omni-Realtime API. - [Server events](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/server-events.md): Server events for the Qwen-Omni-Realtime API, including tool calling (function calling) events. - [Python SDK](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/omni-realtime-python-sdk.md): The key interfaces and request parameters for Qwen-Omni real-time using the DashScope Python SDK. - [Java SDK](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/omni-realtime-java-sdk.md): The key interfaces and request parameters for Qwen-Omni real-time DashScope Java SDK. - [Real-time multimodal interaction flow](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/omni-realtime-interaction-process.md): This topic describes the real-time multimodal interaction flow between the server and the client. - [Voice cloning API reference](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/qwen-omni-voice-cloning.md): Clone a voice from 10-20 seconds of audio without training. This document covers the voice cloning API parameters and usage. For model invocation, see Qwen-Omni-Realtime or Non-real-time (Qwen-Omni) . ## Realtime API - [Realtime API overview](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/realtime-api-overview.md): The Realtime API provides multiple transport protocols, each optimized for different requirements such as performance, latency, weak-network resilience, and integration cost. - [SDK download](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/realtime-sdk-download.md): Download the AOQ SDK for Android, iOS, HarmonyOS, Windows, macOS, Electron, and Linux, and view release notes for each version. - [Token authentication](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/realtime-token-authentication.md): Learn how the Realtime API authenticates connections with tokens, including how to get an API key and how to authenticate over the WebSocket, WebRTC, and AOQ protocols. - [Connecting to models and applications](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/realtime-connect-model.md): Connect to Realtime API models and applications over the AOQ, WebRTC, and WebSocket protocols. This topic covers the connection flow, sequence diagrams, and code examples for each protocol. - [AOQ SDK overview](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/realtime-api-aoq-sdk-desc.md): The AOQ SDK targets real-time multimodal scenarios and helps developers quickly build real-time interactive applications on the Alibaba Cloud Realtime API. - [AOQ Client SDK Android API reference](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/aoq-android-sdk-reference.md): The AOQ Client SDK for Android provides a full set of real-time audio/video communication capabilities, including engine lifecycle management, audio/video capture and playback, codec configuration, external audio stream injection, audio file mixing, real-time messaging, and audio/video frame callbacks. This document is the complete Java API reference for the Android platform. - [AOQ Client SDK iOS API reference](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/aoq-ios-sdk-reference.md): AOQ Client SDK iOS API reference, covering engine lifecycle, audio/video device management, codec configuration, media stream control, audio file playback, external audio streams, real-time messaging, frame callbacks, delegate protocols, and data types and enumerations. - [AOQ Client SDK HarmonyOS API reference](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/aoq-harmony-sdk-reference.md): AOQ Client SDK HarmonyOS (OHOS) API reference. The platform language is ArkTS (.ets), bridged to the C++ engine via NAPI. - [AOQ Client SDK Windows API reference](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/aoq-windows-sdk-reference.md): This topic describes the C++ APIs, callbacks, and data types of AOQ Client SDK for Windows. - [AOQ Client SDK macOS API reference](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/aoq-macos-sdk-reference.md): This topic describes the Objective-C APIs, callbacks, and data types of AOQ Client SDK for macOS. - [AOQ Client SDK Electron API reference](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/aoq-electron-sdk-reference.md): This topic describes the TypeScript APIs, events, and data types of AOQ Client SDK for Electron. The SDK supports macOS x64/arm64 and Windows x64 and requires Node.js 16 or later. - [AOQ Client SDK Linux Python API reference](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/aoq-linux-python-sdk-reference.md): This topic describes the Python APIs, callbacks, and data types of AOQ Client SDK for Linux. - [AOQ Client SDK Linux C++ API reference](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/aoq-linux-cpp-sdk-reference.md): This topic describes the C++ APIs, callbacks, and data types of AOQ Client SDK for Linux. - [Connection state management](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/aoq-connection-management.md): Describes the connection state machine of the AOQ Client SDK and the corresponding API calls. - [Media stream sending control](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/aoq-media-stream-control.md): enableSendMediaStream controls whether the client sends audio or video media streams to the AI service, giving you precise control over when media transmission starts in AOQ protocol scenarios. - [Common audio features](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/aoq-audio-features.md): The AOQ Client SDK provides comprehensive audio capabilities, covering audio capture, playback, codec configuration, speaker management, file mixing, external audio stream injection, and audio frame data callbacks. This document introduces common audio features across Android (Java), iOS (Objective-C), and HarmonyOS (ArkTS). - [Custom audio playback](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/aoq-custom-audio-playback.md): The AOQ Client SDK supports custom audio playback. Through the audio frame callback mechanism, decoded PCM data is delivered to the application layer, letting you implement your own audio rendering logic. - [Custom audio capture](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/aoq-custom-audio-capture.md): Describes how to implement custom audio capture with the AOQ Client SDK, including adding external audio streams, pushing PCM data, and managing stream lifecycle. - [Common video features](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/aoq-video-features.md): The AOQ Client SDK provides comprehensive video capabilities, covering video capture, rendering, codec configuration, frame data callbacks, and external video input. This document introduces common video features across Android (Java), iOS (Objective-C), and HarmonyOS (ArkTS). - [Custom video input](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/aoq-custom-video-input.md): Describes the two custom video input modes supported by the AOQ Client SDK — raw frame mode and encoded frame mode — including configuration and code examples for each. ## Text Embedding - [Synchronous API](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/text-embedding-synchronous-api.md): The general-purpose text embedding model converts text data into numerical vectors for downstream tasks like semantic search, recommendation, clustering, and classification. - [Multimodal embedding API](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/multimodal-embedding-api-reference.md): Multimodal embedding models convert text, images, and videos into embeddings in a shared semantic space to enable cross-modal retrieval, content classification, and similarity search. - [Text rerank](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/text-rerank-api.md): A rerank model re-scores documents returned by initial retrieval, surfacing the most relevant results at the top. ## More Models - [Intent recognition](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/intent-detect-capability.md): Identifies user intents in milliseconds and selects appropriate tools to address queries. - [Qwen-MT API reference](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/qwen-mt-api.md): Input and output parameters for calling Qwen-MT through the OpenAI compatible interface or the DashScope API. - [Qwen-Deep-Research API reference](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/qwen-deep-research-api.md): Request and response parameters for calling the Qwen-Deep-Research model through the DashScope API. - [Qwen-OCR API reference](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/qwen-vl-ocr-api-reference.md): Extract text, structured data, and key information from images using the Qwen-OCR model. Qwen-OCR supports two API protocols: the OpenAI-compatible API and the DashScope API. ## Toolkit/Framework - [OpenAI compatible - Chat](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/compatibility-of-openai-with-dashscope.md): The Qwen models on Model Studio support OpenAI compatible interfaces. You can migrate your existing OpenAI code to Model Studio by changing only the API key, base URL, and model name. - [OpenAI-compatible - Responses](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/compatibility-with-openai-responses-api.md): Alibaba Cloud Model Studio supports the OpenAI-compatible Responses API. Building on the Chat Completions API, the Responses API streamlines native agent functionality. - [Completions API](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/completions.md): The Completions API is designed for text completion scenarios, such as code completion and content continuation. - [OpenAI-compatible - Vision](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/qwen-vl-compatible-with-openai.md): Qwen vision models on Alibaba Cloud Model Studio are compatible with the OpenAI interface specification. You only need to modify three parameters to migrate your existing OpenAI applications to Model Studio: - [OpenAI compatible file interface](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/openai-file-interface.md): Upload files for document Q&A and data extraction with Qwen-Long and Qwen-Doc-Turbo . This interface also supports uploading input files for batch tasks. - [OpenAI-compatible - Batch (file input)](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/batch-interfaces-compatible-with-openai.md): Alibaba Cloud Model Studio provides an OpenAI-compatible Batch File API. Submit requests in bulk through files. The system processes them asynchronously and returns results when all requests complete or the maximum wait time is reached. Costs are only 50% of real-time calls. Ideal for data analytics, model evaluation, and other large-scale workloads where latency is not critical. - [OpenAI-compatible - Batch Chat](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/openai-compatible-batch-chat.md): For non-real-time scenarios like data annotation and content generation, the Batch Chat API offers a low-cost, high-concurrency alternative using the same synchronous call method. Limited-time 50% discount available. - [OpenAI compatible - Embedding](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/embedding-interfaces-compatible-with-openai.md): Alibaba Cloud Model Studio's embedding models are compatible with the OpenAI API. To migrate your existing OpenAI applications to Model Studio, you only need to adjust the following three parameters: - [OpenAI-compatible - Conversations](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/openai-compatible-conversations.md): Manually managing message lists for conversations that span multiple devices or have long interruptions can lead to context loss. Alibaba Cloud Model Studio provides an OpenAI-compatible Conversations API that you can use with the Responses API to automatically inject historical context. This eliminates the need for manual message synchronization and ensures conversational continuity across different scenarios and devices. ## Model Production - [Model tuning](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/fine-tuning-jobs-api.md): Create custom models by fine-tuning. - [Text Generation - Create a Tuning Job](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/create-fine-tuning-job-api.md): Create a model fine-tuning training job for text generation. Datasets can be uploaded via API or mounted from OSS. - [Model import](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/custom-models-api.md): Import fine-tuned model files from OSS into Model Studio. The API supports creating, querying, listing, and deleting import tasks. - [Model compression](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/model-compression-api.md): Compress models using techniques like quantization to reduce inference costs. - [Image Generation - Create a Tuning Job](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/image-generation-create-fine-tuning-job-api.md): Create a fine-tuning job for image generation models. Datasets can be uploaded via the API or mounted from OSS. - [Video Generation - Create a Tuning Job](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/video-generation-create-fine-tuning-job-api.md): Create a model fine-tuning training job for video generation. Datasets can be uploaded via API or mounted from OSS. - [Query fine-tuning task details](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/get-fine-tuning-job-api.md): Query the details of a specific model fine-tuning task. - [Checkpoint management](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/list-checkpoints-api.md): Manage checkpoints generated by fine-tune jobs: list, export, query validation results, and Checkpoint object reference. - [Model deployment](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/deployments-api.md): Deploy your fine-tuned or imported models as an online inference service. - [Create deployment](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/create-deployment-api.md): Create a model deployment task. - [Image Generation - Deploy Model](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/image-generation-deploy-model-api.md): Publish a trained image generation model as an online API service. - [Video Generation - Deploy Model](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/video-generation-deploy-model-api.md): Publish a trained model as an online API service. - [Get deployment details](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/get-deployment-api.md): Model deployment management API, applicable to all model types including text, image, video, and speech . Supports querying deployment status and list, modifying throttling, scaling, and deleting deployments. ## File Management - [File Management](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/file-management-api.md): Manage your files on the Model Studio platform. You can upload, query, list, and delete them. - [Upload a file](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/upload-file-api.md): Upload files to the Model Studio platform. You can upload multiple files at once and reuse them across different jobs. - [Retrieve file details](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/get-file-api.md): Retrieves the details of a file by its file ID. ## More - [Generate a temporary API key](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/generate-temporary-api-key.md): To call model services from untrusted environments such as browsers and mobile apps, use a secure backend service to generate temporary API keys. This prevents your permanent API key from being exposed. - [Asynchronous task management API](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/manage-asynchronous-tasks.md): Some Model Studio models (like image and video generation) use asynchronous invocation due to long processing times. The typical workflow is: create a task to get an ID, then query the result using that ID. Model Studio provides general-purpose task APIs to query individual results, check multiple task statuses in batch, and cancel queued tasks. - [Receive task completion notifications via HTTP callback or MQ](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/async-task-api.md): Instead of polling for task results, you can configure an HTTP callback URL or RocketMQ to receive task completion notifications from EventBridge. Once notified, query the result API once to get the output. - [Model calls in a sub-workspace](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/model-calling-in-sub-workspace.md): Using Qwen-Plus as an example, this topic explains how to invoke a model via an API in a sub-workspace (a non-default workspace). - [Upload local files to get temporary URLs](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/get-temporary-file-url.md): Multimodal models such as Qwen-VL require file URLs for image, video, and audio inputs. Model Studio provides free temporary storage if you lack public URLs: upload a file and get an oss:// URL valid for 48 hours . - [Configure connection reuse for DashScope SDK](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/connection-multiplexing-configuration.md): Without connection reuse, each API call opens a new TCP connection and performs a TLS handshake, adding latency. In high-concurrency scenarios, this overhead causes timeouts and resource waste. Connection reuse eliminates repeated setup, reducing latency and resource consumption. - [List models](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/list-models.md): Call the GET /api/v1/models endpoint to retrieve the list of available models on Model Studio. You can filter by model provider, modality type, model capability, and deployment mode, and get information such as pricing and context length. - [List model quotas](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/list-quotas.md): Call the GET /api/v1/quotas endpoint to query the rate limit quotas for each model under the current API key, including request rate limits (QPS/RPM) and usage limits (TPM), to understand and plan your API usage. - [Update model rate limits](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/update-model-rate-limits.md): Calls the POST /api/v1/models/limits endpoint to update the rate limit quotas (QPM/TPM) of models in a specified workspace. Both merge overlay and delete rate limit operations are supported. - [List model permissions](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/list-model-permissions.md): Call the GET /api/v1/models/permissions endpoint to query the list of authorizable or authorized models and their permission details in the current workspace. - [Update model permissions](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/update-model-permissions.md): Call the POST /api/v1/models/permissions endpoint to update model authorization (inference / fine-tuning / deployment) in a specified workspace. Both per-model authorization and "authorize all" modes are supported. ## Application Calling - [Get the App ID and Workspace ID](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/obtain-the-app-id-and-workspace-id.md): When calling Model Studio apps, such as Agent and Workflow , via an API, you must provide credentials to identify the target app and its workspace: - [Application DashScope API reference](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/application-api-reference.md): Input and output parameters for calling Model Studio applications ( Agent , Workflow ) via the DashScope API, with examples for typical scenarios. ## Application components - [API overview](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/api-bailian-2023-12-29-overview.md) - [Endpoints](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/api-bailian-2023-12-29-endpoint.md) - [RAM authorization](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/api-bailian-2023-12-29-ram.md) - [AddCategory](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/api-bailian-2023-12-29-addcategory.md): Creates a category in a specified workspace to classify and manage files. Each workspace supports a maximum of 500 categories. - [ListCategory](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/api-bailian-2023-12-29-listcategory.md): Retrieves the details of one or more categories in a specified workspace. - [DeleteCategory](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/api-bailian-2023-12-29-deletecategory.md): Permanently deletes a specified category. - [ApplyFileUploadLease](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/api-bailian-2023-12-29-applyfileuploadlease.md): Request an upload lease for uploading knowledge base files or files for agent application conversational interactions. - [AddFile](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/api-bailian-2023-12-29-addfile.md): Imports a file from the temporary storage of Alibaba Cloud Model Studio into an Alibaba Cloud Model Studio data connection (formerly known as application data). - [AddFilesFromAuthorizedOss](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/api-bailian-2023-12-29-addfilesfromauthorizedoss.md): Imports files from an authorized OSS bucket into Alibaba Cloud Model Studio application data. - [ListFile](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/api-bailian-2023-12-29-listfile.md): Retrieves the details of one or more documents in a specified category. - [DescribeFile](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/api-bailian-2023-12-29-describefile.md): Queries the basic information of a file in application data, including the file name, type, and status. - [UpdateFileTag](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/api-bailian-2023-12-29-updatefiletag.md): Updates the tags for a specified file. - [BatchUpdateFileTag](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/api-bailian-2023-12-29-batchupdatefiletag.md): This operation updates document tags in a data connection in batches. - [DeleteFile](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/api-bailian-2023-12-29-deletefile.md): Permanently delete a specified file from application data. Deleting data tables via API is not supported. For details, see the API Guide below. - [DeleteFiles](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/api-bailian-2023-12-29-deletefiles.md): Delete files in batch - [GetParseSettings](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/api-bailian-2023-12-29-getparsesettings.md): Queries the data parsing settings in a specified category. - [GetAvailableParserTypes](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/api-bailian-2023-12-29-getavailableparsertypes.md): Lists all supported parser types based on the input file type (file extension). - [ChangeParseSetting](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/api-bailian-2023-12-29-changeparsesetting.md): Configures the parsing method for specific file types. For example, specifies large model document parsing for .pdf files or Qwen VL parsing for .jpg files. - [AddTable](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/api-bailian-2023-12-29-addtable.md): Add a table for the table data connector. - [UpdateTableFromAuthorizedOss](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/api-bailian-2023-12-29-updatetablefromauthorizedoss.md): Update a table in an Alibaba Cloud Model Studio data connector using a file from an authorized OSS bucket. - [AddConnector](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/api-bailian-2023-12-29-addconnector.md): Creates a connector. This API currently supports only file connectors. - [GetConnector](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/api-bailian-2023-12-29-getconnector.md): Retrieves details about a connector. This operation currently supports only file connectors. - [DeleteConnector](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/api-bailian-2023-12-29-deleteconnector.md): Deletes a connector. - [UpdateConnector](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/api-bailian-2023-12-29-updateconnector.md): Updates a connector. - [CreateIndex](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/api-bailian-2023-12-29-createindex.md): Creates a knowledge base, either an unstructured knowledge base based on documents or audio/video, or a structured knowledge base for data queries or image-based Q&A. - [GetIndexJobStatus](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/api-bailian-2023-12-29-getindexjobstatus.md): Queries the current status of a specified knowledge base creation task or knowledge base document append task. - [SubmitIndexJob](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/api-bailian-2023-12-29-submitindexjob.md): Submits a specified CreateIndex task to complete knowledge base creation. - [SubmitIndexAddDocumentsJob](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/api-bailian-2023-12-29-submitindexadddocumentsjob.md): Appends parsed files to a specified knowledge base. - [Retrieve](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/api-bailian-2023-12-29-retrieve.md): Retrieves relevant text chunks from a knowledge base using vector and keyword search. - [ListIndexDocuments](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/api-bailian-2023-12-29-listindexdocuments.md): Retrieves files and their summary information from a specified knowledge base. - [ListIndexFileDetails](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/api-bailian-2023-12-29-listindexfiledetails.md): Retrieves files and their details from a specified knowledge base. - [DeleteIndexDocument](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/api-bailian-2023-12-29-deleteindexdocument.md): Permanently deletes files from a specified knowledge base. - [ListIndices](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/api-bailian-2023-12-29-listindices.md): Retrieves the list of knowledge bases in a specified workspace. - [DeleteIndex](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/api-bailian-2023-12-29-deleteindex.md): Permanently deletes a specified knowledge base. - [ListChunks](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/api-bailian-2023-12-29-listchunks.md): Queries the list and information of text chunks. - [AddChunk](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/api-bailian-2023-12-29-addchunk.md): Adds chunks to a document search (document), data query (table), or image Q&A (image) knowledge base. - [UpdateChunk](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/api-bailian-2023-12-29-updatechunk.md): Modifies the content and title of a specified text chunk in a knowledge base, and specifies whether the chunk participates in knowledge base retrieval. - [DeleteChunk](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/api-bailian-2023-12-29-deletechunk.md): Deletes specified text chunks from a knowledge base. Deleted text chunks cannot be retrieved or recalled. - [CreatePromptTemplate](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/api-bailian-2023-12-29-createprompttemplate.md): Create a prompt template. - [GetPromptTemplate](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/api-bailian-2023-12-29-getprompttemplate.md): Obtains a prompt template based on the template ID. - [UpdatePromptTemplate](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/api-bailian-2023-12-29-updateprompttemplate.md): Updates a prompt template based on the template ID. - [DeletePromptTemplate](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/api-bailian-2023-12-29-deleteprompttemplate.md): Deletes a prompt template based on the template ID. - [ListPromptTemplates](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/api-bailian-2023-12-29-listprompttemplates.md): Obtains a list of prompt templates. - [ApplyTempStorageLease](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/api-bailian-2023-12-29-applytempstoragelease.md): This interface is intended for pro-code deployment only; other scenarios are currently not supported. It is used to apply for a temporary file upload lease. After obtaining the lease, you must upload the file manually. - [API change records](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/api-bailian-2023-12-29-changeset.md) ## Assistant API (Deprecated) - [Assistants (Deprecated)](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/assistant.md): The assistant API simplifies building assistants, which are a type of Large Language Model (LLM) application. This topic describes the methods provided by the assistant API to manage assistants, such as creating, listing, retrieving, updating, and deleting them. - [Threads (Deprecated)](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/thread.md): This topic describes the Thread class in the assistant API, including how to create, retrieve, modify, and delete threads. - [Messages (Deprecated)](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/message.md): This topic describes how to use the Message class in the assistant API to create, list, retrieve, and modify messages. - [Runs (Deprecated)](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/runs.md): A run represents an execution of an agent in a thread. - [Run steps (Deprecated)](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/run-steps.md): Run steps describe the actions that an agent takes during a run, including model and tool calls. - [Streaming output parameters for the Assistant API (being deprecated)](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/event-streaming.md): The Assistant API streaming output provides real-time results from the Assistant. These results are delivered as an event stream, which contains status information and conversation messages from the Assistant's runtime. To process these messages, you must understand message delta objects and run step delta objects. - [Call examples (Deprecated)](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/call-example.md): This document provides examples that demonstrate how to use agents. ## More - [Service-linked Role](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/bailian-service-linked-role.md): To enable specific features, Alibaba Cloud Model Studio requires a service-linked role (SLR) to access other Alibaba Cloud services, such as Function Compute and Content Moderation, or cloud resources. When you first enable a feature in Alibaba Cloud Model Studio, such as a Function Compute node, the system automatically creates the corresponding service-linked role for you . This topic describes the service-linked roles that are created by Alibaba Cloud Model Studio and explains how to delete them. - [Generate a temporary API key](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/application-obtain-temporary-authentication-token.md): To call model services from untrusted environments such as browsers and mobile apps, use a secure backend service to generate temporary API keys. This prevents your permanent API key from being exposed. - [Knowledge base SearchFilters](https://docs.modelstudio.console.alibabacloud.com/en/model-studio/how-to-use-search-filters.md): The Retrieve API returns text segments from a knowledge base based on semantic similarity. When semantic search returns too many irrelevant results, use SearchFilters to apply metadata-based filtering on top of semantic results. This is effective for structured data such as data tables with well-defined fields.