This document describes how to call the Kimi model inference service deployed on Alibaba Cloud Model Studio.
- OpenAI compatible
- DashScope
- US (Virginia)
- Germany (Frankfurt)
- Singapore
- Japan (Tokyo)
- China (Beijing)
- China (Hong Kong)
base_url for SDK calls is: https://{WorkspaceId}.us-east-1.maas.aliyuncs.com/compatible-mode/v1HTTP request URL: POST https://{WorkspaceId}.us-east-1.maas.aliyuncs.com/compatible-mode/v1/chat/completions{WorkspaceId} with your actual workspace ID.
Prerequisites: You must get an API key and set it as an environment variable. If you use the SDK, you must install the SDK.
Get started
The following examples use text-only input. For multimodal examples, see multimodal call.
- OpenAI compatible
- DashScope
- Anthropic compatible
- Python
- Node.js
- HTTP
Response
Multimodal calls
The kimi-k2.7-code, kimi-k2.6, and kimi-k2.5 models can simultaneously process text, images, or video. Use the enable_thinking parameter to enable thinking mode. The following examples show how to use this capability.
Enable or disable thinking mode
kimi-k2.6 and kimi-k2.5 are hybrid thinking models. These models can reply after thinking or reply directly. You can use the enable_thinking parameter to control whether to enable the thinking mode:
true: Enable thinking modefalse(default): Disables the thinking mode
kimi-k2.7-code is a thinking-only model: thinking mode is always enabled (enable_thinking defaults to true and cannot be disabled), and preserve_thinking defaults to true.
kimi-k2.6 supports passing the thinking process in multi-turn conversations by using the preserve_thinking parameter. For more information, see Pass the thinking process.
The following examples show how to use an image URL and enable thinking mode. The main example demonstrates single-image input, while the commented-out code is an example of multi-image input.
- OpenAI compatible
- DashScope
Video understanding
- Video file
- Image list
-
fps: Controls the frame extraction frequency. The interval between extracted frames is f p s 1 seconds. The value must be in the range of [0.1, 10]. The default value is 2.0.
- For high-motion scenes: Set a higher fps value to capture more detail.
- For static or long videos: Set a lower fps value to improve processing efficiency.
- max_frames: Specifies the maximum number of frames to extract from a video. The default and maximum value is 2000. If the number of frames calculated from the fps value exceeds this limit, the system automatically extracts frames uniformly to stay within the max_frames limit. This parameter is available only when you use the DashScope SDK.
- OpenAI compatible
- DashScope
When passing a video file to the model using the OpenAI SDK or an HTTP request, set the"type"parameter in the user message to"video_url".
Pass a local file
The following examples show how to pass a local file. The OpenAI-compatible API supports only Base64 encoding, while DashScope supports both Base64 encoding and file paths.
- OpenAI compatible
- DashScope
File limitations
- Image limitations
- Video limitations
-
Image resolution:
- Minimum size: Width and height must each exceed
10pixels. - Aspect ratio: The ratio of the longest side to the shortest side must not exceed
200:1. - Maximum resolution: The recommended maximum is
8K(7680x4320). Higher resolutions may cause API call timeouts due to large file sizes or slow network transfers.
- Minimum size: Width and height must each exceed
-
Supported image formats
-
The following formats are supported for resolutions below 4K
(3840x2160):Image format
File extension
MIME type
BMP
.bmp
image/bmp
JPEG
.jpe, .jpeg, .jpg
image/jpeg
PNG
.png
image/png
TIFF
.tif, .tiff
image/tiff
WEBP
.webp
image/webp
HEIC
.heic
image/heic
-
For resolutions between
4K(3840x2160)and8K(7680x4320), only JPEG, JPG, and PNG are supported.
-
The following formats are supported for resolutions below 4K
-
Image size:
- When providing an image via a public URL or local path, its size must not exceed
10 MB. - When using Base64 encoding, the encoded string must not exceed
10 MB.
To compress a file, see How to compress an image or video to meet the size limit.
- When providing an image via a public URL or local path, its size must not exceed
- Number of supported images: When providing multiple images, the total number of tokens for all images and text must not exceed the model's maximum input limit.
Other features
Model | |||||||
|---|---|---|---|---|---|---|---|
kimi-k3 | Supported | Supported | Supported | Supported | Supported | Supported | Supported |
kimi-k2.7-code | Supported | Supported | Supported | Not supported | Not supported | Not supported | Supported |
kimi-k2.6 | Supported | Supported | Supported | Not supported | Not supported | Not supported | Supported |
kimi-k2.5 | Supported | Supported | Supported | Not supported | Not supported | Not supported | Supported |
kimi-k2-thinking | Supported | Supported | Supported | Supported | Not supported | Not supported | Supported |
Moonshot-Kimi-K2-Instruct | Supported | Not supported | Supported | Not supported | Supported | Not supported | Supported |
Dynamically Loaded Tools (Kimi K3)
When an application needs to mount a large number of tools, putting every tool declaration into the request's top-level tools field at once leads to Tool Definition Bloat: every request must carry the description and parameter schema of all tools, driving up token consumption; and the more candidate tools there are, the more likely the model is to pick the wrong tool and construct incorrect call arguments.
Dynamically Loaded Tools let you inject tools on demand during a conversation: mount only a few core tools first, and when the conversation reaches a point where a specific tool is needed, dynamically insert it into messages, thereby reducing token consumption and improving tool-selection accuracy.
tokenization failed error.Inject tool declarations in messages
Insert a message with role set to system into messages, and declare the tools to load via that message's tools field. The format is identical to that of the request's top-level tools field, and you must provide the complete tool information (name, description, parameters).
- A
systemmessage carryingtoolshas the same status as an ordinary message: the tools become visible to the model starting from the position where that message appears in themessageslist. - Dynamically loaded tools coexist with the global tools declared in the request's top-level
toolsfield; the model can see both kinds of tools at the same time. - A dynamically injected tool declaration must be a complete tool definition; you cannot pass only a tool name or reference a globally declared tool.
- A
systemmessage carryingtoolsmust not also carry acontentfield, otherwise the request fails with a 400 error. When using the OpenAI SDK, you can pass thetoolsfield through directly inmessages.
On-demand loading combined with a search tool
There is no dedicated tool-search API. When there are many tools, you can combine a custom search tool with dynamically loaded tools to load tools on demand:
- At the start of the session, declare in the top-level
toolsonly asearch_toolstool implemented by your application backend (it returns matching tool names and summaries by keyword), plus a few core tools that may be used every turn. - Declare the searchable keywords (such as a tool catalog or domain tags) in the system prompt to guide the model to call
search_toolsfirst when it needs a tool. You can settool_choice: "required"on the first request to force the model to search before answering, then restoretool_choiceto"auto"after the search. Changingtool_choicedoes not break the prefix cache. - Based on the results returned by
search_tools, the application dynamically inserts the complete declarations of the corresponding tools intomessagesvia asystemmessage carryingtools. - The model can then call these newly loaded tools directly in subsequent generation.
Notes
- Dynamic tool declarations take effect per request and are not remembered by the server. Whether to keep carrying them in the next request is up to the integrator: keep carrying them and the tools remain available (which also helps hit the prefix cache); stop carrying them and the declaration expires — if the tool is not declared elsewhere, the model cannot call it, and the prefix cache after the change point may miss.
- Appending a dynamic tool declaration at the end of
messagesdoes not affect the cache of the existing prefix; deleting or modifying earlier tool declarations may affect cache hits after the change point. Declaring global tools in the request's top-leveltoolsfield also does not affect cache hits. - A
systemmessage carryingtoolsalso consumes context length, so inject only the tools truly needed by the current conversation. - Dynamic tool declarations use exactly the same format as global
toolsdeclarations, so integrators do not need to maintain two schemas.
Default parameters
Model | enable_thinking | temperature | top_p | presence_penalty | fps | max_frames |
|---|---|---|---|---|---|---|
kimi-k3 | true (thinking mode only, cannot be disabled) | 1.0 | 0.95 | 0.0 | - | - |
kimi-k2.7-code | true (thinking mode only) | 1.0 | 0.95 | 0.0 | 2 | 2000 |
kimi-k2.6 | false | thinking mode: 1.0 non-thinking mode: 0.6 | Both modes: 0.95 | Both modes: 0.0 | 2 | 2000 |
kimi-k2.5 | false | thinking mode: 1.0 non-thinking mode: 0.6 | Both modes: 0.95 | Both modes: 0.0 | 2 | 2000 |
kimi-k2-thinking | - | 1.0 | - | - | - | - |
Moonshot-Kimi-K2-Instruct | - | 0.6 | 1.0 | 0 | - | - |
Models and billing
The Kimi series are large language models from Moonshot AI.
- kimi-k3: Kimi's most capable flagship model to date. It always reasons and uses preserved thinking (thinking-only mode). Supports text and image input (video input is not supported), conversation and agent tasks, and dynamic tool loading.
- kimi-k2.7-code: The most capable Kimi model for coding. It follows long-context instructions more reliably and achieves higher success rates on programming tasks. Supports text, image, and video input, thinking mode, conversation, and agent tasks.
- kimi-k2.6: The newest and most capable model in the Kimi series. It offers improved performance in long-horizon coding, instruction following, and self-correction. Supports text, image, and video input, thinking and non-thinking modes, conversation, and agent tasks.
- kimi-k2.5: It achieves state-of-the-art (SOTA) performance on open-source benchmarks for agent tasks, code generation, visual understanding, and other general intelligence tasks. Supports image, video, and text input, thinking and non-thinking modes, conversation, and agent tasks.
- kimi-k2-thinking: Supports deep thinking mode only. It exposes the reasoning process through the
reasoning_contentfield. It excels at coding and tool calling, and is suitable for use cases that require logical analysis, planning, or deep understanding. - Moonshot-Kimi-K2-Instruct: Does not support deep thinking. It generates responses with lower latency, and is suitable for use cases that need fast, direct answers.
thinking_budget parameter. You cannot use this parameter to limit the thinking length.kimi-k3 does not support the OpenAI-compatible Responses API yet (coming soon). Use the OpenAI-compatible Chat Completions API instead.For pricing, see model invocation billing.For pricing and context window details, see the Model Studio console. Billing is based on input and output token counts.
In thinking mode, the chain of thought counts as output tokens.