Skip to main content
Get Started

Rate limiting

Alibaba Cloud Model Studio applies rate limiting to model calls at the Alibaba Cloud account level, aggregating usage across all RAM users, workspaces, and API keys under the account. Requests are rejected when the limit is exceeded and typically recover automatically within one minute.

Rate limiting rules

  • Account-level rate limiting: Rate limits are applied at the root account level. The usage of all RAM users, workspaces, and API keys under the account is combined.
  • Model-specific rate limiting: Each model has its own rate limit. For more information, see the tables below.

FAQ

Why is rate limiting triggered?

You can identify the type of rate limit triggered based on the error message:
  • Requests rate limit exceeded or You exceeded your current requests list: This indicates that the requests per minute (RPM) limit was triggered.
  • Allocated quota exceeded or You exceeded your current quota: This indicates that the tokens per minute (TPM) limit was triggered.
  • Request rate increased too quickly: The request frequency surged in a short period, triggering system stability protection. This can occur even if the total number of calls has not reached the RPM or TPM limits.
  • For other errors, see Error codes to confirm the cause.
In addition to RPM and TPM, rate limiting may be enforced at the per-second level for requests per second (RPS), which is RPM/60, and tokens per second (TPS), which is TPM/60. Even if the total number of calls per minute does not exceed the limit, a burst of requests in a short time can still trigger rate limiting.

How to view model usage

One hour after you call a model, go to the Monitoring (Singapore or Beijing) page. Set the query conditions, such as the time range and workspace. Then, in the Models area, find the target model and click Monitor in the Actions column to view the model's call statistics. For more information, see the Monitoring document.
Data is updated hourly. During peak periods, there may be an hour-level latency.
image

How long does it take to recover from rate limiting?

Recovery usually occurs within one minute. If other errors occur, see Error codes for troubleshooting. Model response speed is not related to whether you use free quota or paid calls. Free quota and paid calls use the same model service infrastructure, so the generation speed for a given model is the same. Free quota only controls whether a request is accepted: once the free quota is exhausted, the request is rejected with a 403 error (if you have enabled stop when the free quota is used up). This does not reduce the generation speed of any request that has already been accepted.

What factors affect model response speed?

  • Model type: lightweight models (such as qwen-flash) generate responses faster than larger models (such as qwen-max).
  • Output length: the more tokens a response contains, the longer the total generation time.
  • Server load: response speed may fluctuate slightly during peak periods.

Does rate limiting reduce the generation speed of accepted requests?

No. Rate limiting (RPM/TPM) only rejects requests that exceed the limit, returning a 429 error. When the free quota is exhausted (if you have enabled stop when the free quota is used up), requests are rejected with a 403 error. Neither case affects the generation speed of requests that have already been accepted: 429 indicates rate limiting, and 403 indicates quota exhaustion — both are rejection mechanisms, not speed throttling.

How to avoid rate limiting

  1. Choose models with higher rate limits: Stable or latest versions have higher rate limits than dated snapshot versions.
  2. Optimize your call strategy
    • Reduce call frequency: If you receive a Requests rate limit exceeded or You exceeded your current requests list error, lower the API call frequency.
    • Reduce token consumption: If you receive an Allocated quota exceeded or You exceeded your current quota error, shorten the input or limit the output length.
    • Smooth the request rate: If you receive a Request rate increased too quickly error, use uniform scheduling, exponential backoff, or a request queue to distribute requests evenly and avoid sudden peaks.
  3. Add a backup model If rate limiting is triggered, you can switch to a backup model to continue generation. This can reduce the probability of failure and increase throughput. The following code automatically retries with qwen-plus-2025-07-14 after a rate limit is triggered for qwen-plus-2025-07-28. You can also choose a model from a different series (such as qwen-flash) as the backup model to further reduce the risk that the primary and backup models are rate limited at the same time.
    import os
    import asyncio
    from openai import AsyncOpenAI, APIStatusError
    
    # Configuration
    API_KEY = os.getenv("DASHSCOPE_API_KEY")
    # Primary model
    MODEL = "qwen-plus-2025-07-28"
    # Backup model
    BACKUP_MODEL = "qwen-plus-2025-07-14"
    # Test question
    QUESTION = "Who are you?"
    # Concurrency setting
    NUM_REQUESTS = 10
    
    client = AsyncOpenAI(
        api_key=API_KEY,
        # When calling, replace {WorkspaceId} with your actual workspace ID.
        base_url="https://{WorkspaceId}.ap-southeast-1.maas.aliyuncs.com/compatible-mode/v1"
    )
    
    async def send_request(model):
        """Sends a single request."""
        try:
            await client.chat.completions.create(
                model=model,
                messages=[{"role": "user", "content": QUESTION}]
            )
            return True
        except APIStatusError as e:
            if e.status_code == 429:
                print(f"[Rate limit triggered] Model {model}")
                return False
            raise
        except Exception as e:
            print(f"[Request failed] Model {model}, Error: {e}")
            return False
    
    async def task(i):
        # Try the primary model.
        if await send_request(MODEL):
            return True
        # If rate limited, try the backup model.
        return await send_request(BACKUP_MODEL)
    
    async def main():
        results = await asyncio.gather(*(task(i) for i in range(NUM_REQUESTS)))
        print(f"Successful requests: {sum(results)}, Failed requests: {len(results) - sum(results)}")
    
    if __name__ == "__main__":
        asyncio.run(main())
    
  4. Split tasks: Long conversations or large documents can consume many tokens quickly. You can split large batch tasks into smaller batches and submit them at different times.
  5. Use batch inference: For tasks that do not require real-time responses, you can use the Batch API. Batch requests are not subject to real-time rate limits, but you must consider queuing and processing time.
  6. Increase rate limits: If the default rate limits are insufficient, you can increase the temporary TPM quota for a model on the Increase Rate Limits page in the Model Studio console. The increase takes effect immediately. For more information, see Increase temporary rate limits.

How to control token usage or costs

Rate limiting only restricts the request rate per unit of time; it does not cap cumulative usage. To control token usage or costs, use the following methods:
  • Set a spending limit and cost alerts: On the Billing card, configure Cost alerts to enable a monthly spending limit and threshold notifications. You are notified when the threshold is reached, which helps you avoid overspending. For more information, see Query bills and manage costs.
  • Enable stop when the free quota is used up: For models that offer a free quota, you can enable stop when the free quota is used up so that calls stop automatically once the free quota is exhausted, which prevents additional charges. For more information, see Free quota.
  • Monitor model usage: Regularly check the token usage of each model to detect abnormal growth in time. See How to view model usage above.

Will I still be rate limited after I top up my account?

Topping up your account does not change the default RPM and TPM rate limits of a model. Rate limits are configured at the Alibaba Cloud account level and are independent of billing. Topping up (pay-as-you-go) only ensures that your account is not suspended for overdue payments; it does not raise your rate limits. To obtain higher rate limits, see the How to avoid rate limiting section in this topic, or apply for an increase on the Increase Rate Limits page in the Model Studio console. For more information, see Increase temporary rate limits.

Increase temporary rate limits

If the default rate limits are insufficient, you can increase a model's temporary TPM quota in the Model Studio console. The increase takes effect immediately and is valid for 30 days. After it expires, the quota automatically reverts to the system default. This feature is currently available in the China (Beijing) and Singapore regions.
  1. Log on to the Model Studio console and go to the Increase Rate Limits page.
  2. Click Increase Temporary Rate Limit in the upper-right corner.
  3. In the dialog box that appears, select a Model and enter the desired value for Token Account Limit (Token/60 seconds). The dialog box displays the current quota and the maximum configurable limit.
  4. Click Confirm. The increased quota takes effect immediately.
After the quota increase takes effect, you can confirm it in the following ways:
  • On the Increase Rate Limits page, view the models with increased quotas and their corresponding rate limit data in the list.
  • In the Model List, go to the details page of the corresponding model to view the updated rate limit data.
  • The models for which you can temporarily increase quotas are listed in the dialog box on the Increase Rate Limits page.
  • Submitting another request for a model that already has an increased quota is considered a new application, and the validity period is reset to 30 days.
  • Request a quota based on your actual needs. If the provisioned capacity significantly exceeds actual usage for a long time, the system may restore it to the default value after prior notification.
  • If the available rate limit quota still does not meet your needs, contact your account manager to request a further increase.

Text generation - Qwen

Qwen language model

  • Singapore
  • US (Virginia)
  • China (Beijing)
  • Germany (Frankfurt)
  • Hong Kong (China)
  • Japan (Tokyo)
Model nameService deployment scopeRate limiting conditions (triggered if any value is exceeded)
The following limits are per minute. The service may also enforce limits based on requests per second (RPS = RPM/60) and tokens per second (TPS = TPM/60).
Requests per minute (RPM)Tokens per minute (TPM)
Includes input and output tokens.
qwen3.8-maxInternational15,0002,000,000
qwen3.8-flashInternational15,0002,000,000
qwen3.7-maxInternational6001,000,000
qwen3.7-max-2026-06-08International601,000,000
qwen3.7-max-2026-05-20International601,000,000
qwen3.7-max-previewInternational6001,000,000
qwen3.7-max-2026-05-17International6001,000,000
qwen3.6-max-previewInternational6001,000,000
qwen3-maxInternational6001,000,000
qwen3-max-2026-01-23International6001,000,000
qwen3-max-2025-09-23International60100,000
qwen3-max-previewInternational6001,000,000
qwen-max
Rate limiting does not apply to service calls made using the Batch API.
International6001,000,000
qwen3.7-plusInternational15,0005,000,000
qwen3.7-plus-2026-05-26International601,000,000
qwen3.6-plusInternational15,0005,000,000
qwen3.6-plus-2026-04-02International601,000,000
qwen3.7-flashInternational15,0005,000,000
qwen3.7-flash-2026-07-15International15,0005,000,000
qwen3.6-flashInternational15,0005,000,000
qwen3.6-flash-2026-04-16International601,000,000
qwen3.5-plusInternational15,0005,000,000
qwen3.5-plus-2026-04-20International6001,000,000
qwen3.5-plus-2026-02-15International601,000,000
qwen-plus
Rate limiting does not apply to service calls made using the Batch API.
International6001,000,000
qwen-plus-latestInternational6001,000,000
qwen-plus-2025-12-01International1201,000,000
qwen-plus-2025-09-11International1201,000,000
qwen-plus-2025-07-28International60100,000
qwen-plus-2025-07-14(qwen-plus-0714)International60100,000
qwen-plus-2025-04-28(qwen-plus-0428)International601,000,000
qwen-plus-2025-01-25(qwen-plus-0125)International60100,000
qwen3.5-flashInternational15,0005,000,000
qwen3.5-flash-2026-02-23International601,000,000
qwen-flash
Rate limiting does not apply to service calls made using the Batch API.
International6005,000,000
qwen-flash-2025-07-28International6005,000,000
qwq-plusInternational60100,000
qwen-turbo
Rate limiting does not apply to service calls made using the Batch API.
International6005,000,000

Qwen-VL (visual understanding/image-to-text)

  • Singapore
  • US (Virginia)
  • China (Beijing)
  • Germany (Frankfurt)
  • Hong Kong (China)
Model nameService deployment scopeRate limiting conditions (triggered if any value is exceeded)
The following limits are per minute. The service may also enforce limits based on requests per second (RPS = RPM/60) and tokens per second (TPS = TPM/60).
Requests per minute (RPM)Tokens per minute (TPM)
Includes input and output tokens.
qwen3-vl-plusInternational1,2001,000,000
qwen3-vl-plus-2025-12-19International60100,000
qwen3-vl-plus-2025-09-23International1201,000,000
qwen3-vl-flashInternational1,2001,000,000
qwen3-vl-flash-2026-01-22International60100,000
qwen3-vl-flash-2025-10-15International1201,000,000
qwen-vl-maxInternational1,2001,000,000
qwen-vl-plusInternational1,2001,000,000
qvq-maxInternational60100,000

Qwen-Omni (omni-modal)

  • Singapore
  • China (Beijing)
Model nameService deployment scopeRate limiting conditions (triggered if any value is exceeded)
The following limits are per minute. The service may also enforce limits based on requests per second (RPS = RPM/60) and tokens per second (TPS = TPM/60).
Requests per minute (RPM)Tokens per minute (TPM)
Includes input and output tokens.
qwen3.5-omni-flashInternational60100,000
qwen3.5-omni-flash-2026-03-15International60100,000
qwen3.5-omni-plusInternational60100,000
qwen3.5-omni-plus-2026-03-15International60100,000
qwen3-omni-flashInternational60100,000
qwen3-omni-flash-2025-12-01International60100,000
qwen3-omni-flash-2025-09-15International60100,000
qwen-omni-turboInternational60100,000
qwen-omni-turbo-latestInternational60100,000
qwen-omni-turbo-2025-03-26International60100,000

Qwen-Omni-Realtime (real-time multimodal)

  • Singapore
  • China (Beijing)
Model nameService deployment scopeRate limiting conditions (triggered if any value is exceeded)
The following limits are per minute. The service may also enforce limits based on requests per second (RPS = RPM/60) and tokens per second (TPS = TPM/60).
Requests per minute (RPM)Tokens per minute (TPM)
Includes input and output tokens.
qwen3.5-omni-plus-realtimeInternational60100,000
qwen3.5-omni-plus-realtime-2026-03-15International60100,000
qwen3.5-omni-flash-realtimeInternational60100,000
qwen3.5-omni-flash-realtime-2026-03-15International60100,000
qwen3-omni-flash-realtimeInternational60100,000
qwen3-omni-flash-realtime-2025-12-01International60100,000
qwen3-omni-flash-realtime-2025-09-15International60100,000
qwen-omni-turbo-realtimeInternational6010,000
qwen-omni-turbo-realtime-latestInternational6010,000
qwen-omni-turbo-realtime-2025-05-08International6010,000

Qwen-OCR (text extraction)

  • Singapore
  • US (Virginia)
  • China (Beijing)
  • Germany (Frankfurt)
Model nameService deployment scopeRate limiting conditions (triggered if any value is exceeded)
The following limits are per minute. The service may also enforce limits based on requests per second (RPS = RPM/60) and tokens per second (TPS = TPM/60).
Requests per minute (RPM)Tokens per minute (TPM)
Includes input and output tokens.
qwen-vl-ocrInternational6006,000,000
qwen-vl-ocr-2025-11-20International1,2006,000,000

Qwen math model

  • China (Beijing)
Model nameRate limiting conditions (triggered if any value is exceeded)
The following limits are per minute. The service may also enforce limits based on requests per second (RPS = RPM/60) and tokens per second (TPS = TPM/60).
Requests per minute (RPM)Tokens per minute (TPM)
Includes input and output tokens
qwen-math-plus1,2001,000,000
qwen-math-plus-latest1,2001,000,000
qwen-math-plus-2024-09-19(qwen-math-plus-0919)60100,000
qwen-math-plus-2024-08-16(qwen-math-plus-0816)1020,000
qwen-math-turbo12001,000,000

Qwen-Coder

  • Singapore
  • US (Virginia)
  • China (Beijing)
  • Germany (Frankfurt)
Model nameService deployment scopeRate limiting conditions (triggered if any value is exceeded)
The following limits are per minute. The service may also enforce limits based on requests per second (RPS = RPM/60) and tokens per second (TPS = TPM/60).
Requests per minute (RPM)Tokens per minute (TPM)
Includes input and output tokens.
qwen3-coder-plusInternational2,4002,000,000
qwen3-coder-plus-2025-09-23International6001,000,000
qwen3-coder-plus-2025-07-22International601,000,000
qwen3-coder-flashInternational6005,000,000
qwen3-coder-flash-2025-07-28International6005,000,000

Qwen translation model

  • Singapore
  • US (Virginia)
  • China (Beijing)
  • Germany (Frankfurt)
Model nameService deployment scopeRate limiting conditions (triggered if any value is exceeded)
The following limits are per minute. The service may also enforce limits based on requests per second (RPS = RPM/60) and tokens per second (TPS = TPM/60).
Requests per minute (RPM)Tokens per minute (TPM)
Includes input and output tokens.
qwen-mt-plusInternational60100,000
qwen-mt-flashInternational60100,000
qwen-mt-liteInternational60100,000
qwen-mt-turboInternational60100,000

Qwen data mining model

  • China (Beijing)
Model nameRate limiting conditions (triggered if any value is exceeded)
The following limits are per minute. The service may also enforce limits based on requests per second (RPS = RPM/60) and tokens per second (TPS = TPM/60).
Requests per minute (RPM)Tokens per minute (TPM)
Includes input and output tokens
qwen-doc-turbo6003,000,000

Qwen deep research model

  • China (Beijing)
Model nameRate limiting conditions (triggered if any value is exceeded)
The following limits are per minute. The service may also enforce limits based on requests per second (RPS = RPM/60) and tokens per second (TPS = TPM/60).
Requests per minute (RPM)Tokens per minute (TPM)
Includes input and output tokens
qwen-deep-research1201,200,000

Text generation - Qwen - Open source

Qwen language model open source

  • Singapore
  • US (Virginia)
  • China (Beijing)
  • Germany (Frankfurt)
Model nameService deployment scopeRate limiting conditions (triggered if any value is exceeded)
The following limits are per minute. The service may also enforce limits based on requests per second (RPS = RPM/60) and tokens per second (TPS = TPM/60).
Requests per minute (RPM)Tokens per minute (TPM)
Includes input and output tokens.
qwen3.8-2.4t-a95bInternational5,0005,000,000
qwen3.8-27bInternational5,0005,000,000
qwen3.6-35b-a3bInternational6001,000,000
qwen3.6-27bInternational6001,000,000
qwen3.5-397b-a17bInternational6001,000,000
qwen3.5-122b-a10bInternational6001,000,000
qwen3.5-27bInternational6001,000,000
qwen3.5-35b-a3bInternational6005,000,000
qwen3-next-80b-a3b-thinkingInternational6001,000,000
qwen3-next-80b-a3b-instructInternational6001,000,000
qwen3-235b-a22b-thinking-2507International6001,000,000
qwen3-235b-a22b-instruct-2507International6001,000,000
qwen3-30b-a3b-thinking-2507International6005,000,000
qwen3-30b-a3b-instruct-2507International6005,000,000
qwen3-235b-a22bInternational6001,000,000
qwen3-32bInternational6001,000,000
qwen3-30b-a3bInternational6001,000,000
qwen3-14bInternational6001,000,000
qwen3-8bInternational6001,000,000

Qwen-VL (visual understanding/image-to-text)

  • Singapore
  • US (Virginia)
  • China (Beijing)
  • Germany (Frankfurt)
Model nameService deployment scopeRate limiting conditions (triggered if any value is exceeded)
The following limits are per minute. The service may also enforce limits based on requests per second (RPS = RPM/60) and tokens per second (TPS = TPM/60).
Requests per minute (RPM)Tokens per minute (TPM)
Includes input and output tokens.
qwen3-vl-32b-thinkingInternational60100,000
qwen3-vl-32b-instructInternational60100,000
qwen3-vl-30b-a3b-thinkingInternational60100,000
qwen3-vl-30b-a3b-instructInternational60100,000
qwen3-vl-8b-thinkingInternational60100,000
qwen3-vl-8b-instructInternational60100,000
qwen3-vl-235b-a22b-thinkingInternational60100,000
qwen3-vl-235b-a22b-instructInternational60100,000

Qwen3-Omni

  • Singapore
  • China (Beijing)
Model nameService deployment scopeRate limiting conditions (triggered if any value is exceeded)
The following limits are per minute. The service may also enforce limits based on requests per second (RPS = RPM/60) and tokens per second (TPS = TPM/60).
Requests per minute (RPM)Tokens per minute (TPM)
Includes input and output tokens.
qwen2.5-omni-7bInternational60100,000

Qwen3-Omni-Captioner

  • Singapore
  • China (Beijing)
Model nameService deployment scopeRate limiting conditions (triggered if any value is exceeded)
The following limits are per minute. The service may also enforce limits based on requests per second (RPS = RPM/60) and tokens per second (TPS = TPM/60).
Requests per minute (RPM)Tokens per minute (TPM)
Includes input and output tokens.
qwen3-omni-30b-a3b-captionerInternational60100,000

Qwen-Coder

  • Singapore
  • US (Virginia)
  • China (Beijing)
  • Germany (Frankfurt)
Model nameService deployment scopeRate limiting conditions (triggered if any value is exceeded)
The following limits are per minute. The service may also enforce limits based on requests per second (RPS = RPM/60) and tokens per second (TPS = TPM/60).
Requests per minute (RPM)Tokens per minute (TPM)
Includes input and output tokens.
qwen3-coder-nextInternational6001,000,000
qwen3-coder-480b-a35b-instructInternational6001,000,000
qwen3-coder-30b-a3b-instructInternational6001,000,000

Text generation - Third-party models

DeepSeek

  • Singapore
  • US (Virginia)
  • China (Beijing)
  • Germany (Frankfurt)
  • Japan (Tokyo)
  • Hong Kong (China)
Model nameService deployment scopeRate limiting conditions (triggered if any value is exceeded)
The following limits are per minute. The service may also enforce limits based on requests per second (RPS = RPM/60) and tokens per second (TPS = TPM/60).
Requests per minute (RPM)Tokens per minute (TPM)
Includes input and output tokens.
deepseek-v4-proInternational10,0001,200,000
deepseek-v4-pro-0813International10,0001,200,000
deepseek-v4-flash-0731International15,0001,200,000
deepseek-v4-flashInternational10,0001,200,000
deepseek-v3.2International10,0001,200,000

Kimi

  • China (Beijing)
  • US (Virginia)
  • Germany (Frankfurt)
  • Hong Kong (China)
  • Japan (Tokyo)
  • Singapore
Model nameRate limiting conditions (triggered if any value is exceeded)
The following limits are per minute. The service may also enforce limits based on requests per second (RPS = RPM/60) and tokens per second (TPS = TPM/60).
Requests per minute (RPM)Tokens per minute (TPM)
Includes input and output tokens
kimi-k315,0001,200,000
kimi-k2.7-code5001,000,000
kimi-k2.65001,000,000
kimi-k2.55001,000,000
kimi-k2-thinking5001,000,000
Moonshot-Kimi-K2-Instruct5001,000,000

MiniMax

  • China (Beijing)
Model nameRate limiting conditions (triggered if any value is exceeded)
The following limits are per minute. The service may also enforce limits based on requests per second (RPS = RPM/60) and tokens per second (TPS = TPM/60).
Requests per minute (RPM)Tokens per minute (TPM)
Includes input and output tokens
MiniMax-M2.55001,000,000

GLM

  • US (Virginia)
  • China (Beijing)
  • Germany (Frankfurt)
  • Singapore
  • Hong Kong (China)
  • Japan (Tokyo)
Model nameService deployment scopeRate limiting conditions (triggered if any value is exceeded)
The following limits are per minute. The service may also enforce limits based on requests per second (RPS = RPM/60) and tokens per second (TPS = TPM/60).
Requests per minute (RPM)Tokens per minute (TPM)
Includes input and output tokens.
glm-5.2Global5001,000,000
glm-5.2-usUS5001,000,000
glm-5.1Global5001,000,000

GLM-Z.AI direct supply

  • Singapore
Model nameRate limits (triggered if any value is exceeded)
The following are the limits per minute. The service may also enforce limits based on requests per second (RPS = RPM/60) and tokens per second (TPS = TPM/60).
Requests per minute (RPM)Tokens per minute (TPM)
Including input and output tokens
ZHIPU/GLM-5.32003,000,000
ZHIPU/GLM-5.22003,000,000

Image generation

Qwen-Image

  • Singapore
  • China (Beijing)
  • Hong Kong (China)
  • Germany (Frankfurt)
  • Japan (Tokyo)

Model name

Service deployment scope

Rate limiting conditions (triggered if any value is exceeded)

Task submission API call limit

Number of concurrent tasks (concurrency)

qwen-image-3.0-pro

International

5 request/minute

No limit for sync APIs / 10 for async APIs

qwen-image-3.0

International

20 request/minute

No limit for sync APIs / 10 for async APIs

qwen-image-2.0-pro

International

2 times/minute

No limit for synchronous APIs

qwen-image-2.0-pro-2026-06-22

International

2 times/minute

No limit for synchronous APIs

qwen-image-2.0-pro-2026-04-22

International

2 times/minute

No limit for synchronous APIs

qwen-image-2.0-pro-2026-03-03

International

2 times/minute

No limit for synchronous APIs

qwen-image-2.0

International

2 times/second

No limit for synchronous APIs

qwen-image-2.0-2026-03-03

International

2 times/second

No limit for synchronous APIs

qwen-image-max

International

2 times/minute

No limit for synchronous APIs

qwen-image-max-2025-12-30

International

2 times/minute

No limit for synchronous APIs

qwen-image-plus

International

2 times/second

No limit for synchronous APIs / 2 for asynchronous APIs

qwen-image-plus-2026-01-09

International

2 times/second

No limit for synchronous APIs

qwen-image

International

2 times/second

No limit for synchronous APIs / 2 for asynchronous APIs

qwen-image-edit-max

International

2 times/minute

No limit for synchronous APIs

qwen-image-edit-max-2026-01-16

International

2 times/minute

No limit for synchronous APIs

qwen-image-edit-plus

International

2 times/second

No limit for synchronous APIs

qwen-image-edit-plus-2025-12-15

International

2 times/second

No limit for synchronous APIs

qwen-image-edit-plus-2025-10-30

International

2 times/second

No limit for synchronous APIs

qwen-image-edit

International

2 times/second

No limit for synchronous APIs

qwen-mt-image-2.0

International

60 times/minute

2

Text-to-image - Z-Image

  • Singapore
  • China (Beijing)

Model name

Service deployment scope

Rate limiting conditions (triggered if any value is exceeded)

RPS limit for task submission API

Number of concurrent tasks (concurrency)

z-image-turbo

International

2

No limit for synchronous APIs

Wanxiang

  • Singapore
  • US (Virginia)
  • China (Beijing)
  • Germany (Frankfurt)

Model name

Service deployment scope

Rate limiting conditions (triggered if any value is exceeded)

RPS limit for task submission API

Number of concurrent tasks (concurrency)

wan2.7-image-pro

International

5

5

wan2.7-image

International

5

5

wan2.6-image

International

5

5

wan2.6-t2i

International

5

5

wan2.5-t2i-preview

International

5

5

wan2.2-t2i-flash

International

2

2

wan2.2-t2i-plus

International

2

2

wan2.1-t2i-turbo

International

2

2

wan2.1-t2i-plus

International

2

2

wan2.5-i2i-preview

International

5

5

OutfitAnyone

  • China (Beijing)

Model name

Rate limiting conditions (triggered if any value is exceeded)

RPS limit for job submission API

Number of concurrent tasks

aitryon-plus

10

5

aitryon-parsing-v1

10

No limit for synchronous APIs

Video generation

HappyHorse series

  • Singapore
  • US (Virginia)
  • China (Beijing)
  • Germany (Frankfurt)
  • Japan (Tokyo)

Model name

Service deployment scope

Rate limiting conditions (triggered if any value is exceeded)

RPS limit for task submission API

Number of concurrent tasks (concurrency)

happyhorse-1.1-t2v

International

5

5

happyhorse-1.1-i2v

International

5

5

happyhorse-1.1-r2v

International

5

5

happyhorse-1.0-t2v

International

5

5

happyhorse-1.0-i2v

International

5

5

happyhorse-1.0-r2v

International

5

5

happyhorse-1.0-video-edit

International

5

5

Wanxiang series

  • Singapore
  • US (Virginia)
  • China (Beijing)
  • Germany (Frankfurt)
  • Germany (Frankfurt)
  • Japan (Tokyo)
  • China (Hong Kong)

Model name

Service deployment scope

Rate limiting conditions (triggered if any value is exceeded)

RPS limit for task submission API

Number of concurrent tasks (concurrency)

wan3.0-video-prime

International

5

5

wan3.0-video

International

5

5

wan2.7-r2v-2026-06-12

Global

5

5

wan2.7-t2v-2026-06-12

International

5

5

wan2.7-t2v-2026-04-25

International

5

5

wan2.7-t2v

International

5

5

wan2.6-t2v

International

5

5

wan2.5-t2v-preview

International

5

5

wan2.2-t2v-plus

International

2

2

wan2.1-t2v-turbo

International

2

2

wan2.1-t2v-plus

International

2

2

wan2.7-i2v-2026-04-25

International

5

5

wan2.7-i2v

International

5

5

wan2.6-i2v-flash

International

5

5

wan2.6-i2v

International

5

5

wan2.5-i2v-preview

International

5

5

wan2.2-i2v-flash

International

2

2

wan2.1-i2v-plus

International

2

2

wan2.1-i2v-turbo

International

2

2

wan2.2-i2v-plus

International

2

2

wan2.2-kf2v-flash

International

2

2

wan2.1-kf2v-plus

International

1

2

wan2.1-vace-plus

International

2

2

wan2.7-videoedit

International

5

5

wan2.7-r2v

International

5

5

wan2.6-r2v-flash

International

5

5

wan2.6-r2v

International

5

5

wan2.2-animate-move

International

5

1

wan2.2-animate-mix

International

5

1

AnimateAnyone

  • China (Beijing)

Model name

RPS limit for task submission API

Number of concurrent tasks

animate-anyone-detect-gen2

5

No limit for synchronous APIs

animate-anyone-template-gen2

5

1

Only one job runs at a time. Other jobs in the queue are in a waiting state.

animate-anyone-gen2

5

1

Only one job runs at a time. Other jobs in the queue are in a waiting state.

EMO

  • China (Beijing)

Model name

RPS limit for task submission API

Number of concurrent tasks

emo-detect-v1

5

No limit for synchronous APIs

emo-v1

5

1

Only one job runs at a time. Other jobs in the queue are in a waiting state.

LivePortrait

  • China (Beijing)

Model name

RPS limit for task submission API

Number of concurrent tasks

liveportrait-detect

5

No limit for synchronous APIs

liveportrait

5

1

Only one job runs at a time. Other jobs in the queue are in a waiting state.

VideoRetalk

  • China (Beijing)

Model name

RPS limit for task submission API

Number of concurrent tasks

videoretalk

1

1

Only one job runs at a time. Other jobs in the queue are in a waiting state.

Emoji

  • China (Beijing)

Model name

RPS limit for task submission API

Number of concurrent tasks

emoji-detect-v1

1

No limit for synchronous APIs

emoji-v1

1

1

Only one job runs at a time. Other jobs in the queue are in a waiting state.

Video style transform

  • China (Beijing)

Model name

RPS limit for task submission API

Number of concurrent tasks

video-style-transform

20

2

Only one job runs at a time. Other jobs in the queue are in a waiting state.

Music generation

  • China (Beijing)

Model name

Requests per minute (RPM)

fun-music-preview

180

fun-music-v1

180

Voice chat

Realtime voice chat

  • Singapore
  • China (Beijing)
Model nameService deployment scopeRate limiting conditions (triggered if any value is exceeded)
The following limits are per minute. The service may also enforce limits based on requests per second (RPS = RPM/60) and tokens per second (TPS = TPM/60).
Requests per minute (RPM)Tokens per minute (TPM)
Including input and output tokens
qwen-audio-3.0-realtime-plusInternational60100,000
qwen-audio-3.0-realtime-flashInternational60100,000

Speech synthesis (text-to-speech)

Qwen-Audio-TTS

  • Singapore
  • China (Beijing)

Model name

Service deployment scope

Job submission API RPS limit

qwen-audio-3.0-tts-plus

International

3

qwen-audio-3.0-tts-flash

International

3

Qwen-TTS

  • Singapore
  • China (Beijing)
  • Qwen3-TTS-Instruct-Flash
  • Qwen3-TTS-VD
  • Qwen3-TTS-VC
  • Qwen3-TTS-Flash

Model name

Service deployment scope

Requests per minute (RPM)

qwen3-tts-instruct-flash

International

180

qwen3-tts-instruct-flash-2026-01-26

International

180

Qwen-TTS-Realtime

  • Singapore
  • China (Beijing)
  • Qwen3-TTS-Instruct-Flash-Realtime
  • Qwen3-TTS-VD-Realtime
  • Qwen3-TTS-VC-Realtime
  • Qwen3-TTS-Flash-Realtime

Model name

Service deployment scope

Requests per minute (RPM)

qwen3-tts-instruct-flash-realtime

International

180

qwen3-tts-instruct-flash-realtime-2026-01-22

International

180

Qwen-TTS voice cloning

  • Singapore
  • China (Beijing)

Model name

Service deployment scope

Requests per minute (RPM)

qwen-voice-enrollment

International

180

Qwen-TTS voice design

  • Singapore
  • China (Beijing)

Model name

Service deployment scope

Requests per minute (RPM)

qwen-voice-design

International

180

CosyVoice

  • Singapore
  • China (Beijing)

Model name

Service deployment scope

Job submission API RPS limit

cosyvoice-v3-plus

International

3

cosyvoice-v3-flash

International

Qwen-Audio-TTS/CosyVoice voice cloning/design

Qwen-Audio-TTS/CosyVoice voice cloning/design models share a single model and a single rate limit quota.
  • Singapore
  • China (Beijing)

Model name

Service deployment scope

Job submission API RPS limit

voice-enrollment

International

10

Speech recognition (speech-to-text) and translation (speech to text in a specified language)

Qwen3-LiveTranslate-Flash

  • Singapore
  • China (Beijing)
Model nameService deployment scopeRate limiting conditions (rate limiting is triggered when any value is exceeded)
The following are per-minute rate limiting conditions. The service may also enforce RPS (RPM/60) and TPS (TPM/60) limits
Requests per minute (RPM)Tokens per minute (TPM)
Including input and output tokens
qwen3-livetranslate-flashInternational100100,000
qwen3-livetranslate-flash-2025-12-01International100100,000

Qwen-LiveTranslate-Flash-Realtime

  • Singapore
  • China (Beijing)
Model nameService deployment scopeRate limiting conditions (rate limiting is triggered when any value is exceeded)
The following are per-minute rate limiting conditions. The service may also enforce RPS (RPM/60) and TPS (TPM/60) limits
Requests per minute (RPM)Tokens per minute (TPM)
Including input and output tokens
qwen3.5-livetranslate-flash-realtimeInternational10100,000
qwen3.5-livetranslate-flash-realtime-2026-05-19International
qwen3-livetranslate-flash-realtimeInternational
qwen3-livetranslate-flash-realtime-2025-09-22International

Qwen-Audio-3.0-ASR-Flash-Streaming

  • Singapore
  • China (Beijing)

Model name

Service deployment scope

Requests per second (RPS)

qwen-audio-3.0-asr-flash-streaming

International

20

Qwen-Audio-3.0-ASR-Flash-Filetrans

  • Singapore
  • China (Beijing)

Model name

Service deployment scope

Requests per minute (RPM)

qwen-audio-3.0-asr-flash-filetrans

International

600

Qwen-Audio-3.0-ASR-Flash

  • Singapore
  • China (Beijing)

Model name

Service deployment scope

Requests per minute (RPM)

qwen-audio-3.0-asr-flash

International

600

Qwen-ASR

  • Singapore
  • US (Virginia)
  • China (Beijing)
  • Qwen3-ASR-Flash-Filetrans
  • Qwen3-ASR-Flash

Model name

Service deployment scope

Requests per minute (RPM)

qwen3-asr-flash-filetrans

International

100

qwen3-asr-flash-filetrans-2025-11-17

International

Qwen-ASR-Realtime

  • Singapore
  • China (Beijing)

Model name

Service deployment scope

Requests per second (RPS)

qwen3-asr-flash-realtime

International

20

qwen3-asr-flash-realtime-2026-02-10

International

qwen3-asr-flash-realtime-2025-10-27

International

Paraformer

  • China (Beijing)

Model name

Job submission API RPS limit

paraformer-realtime-v2

20

paraformer-realtime-8k-v2

Model name

Requests per minute (RPM)

paraformer-v2

1,200

Model name

Job submission API RPS limit

Concurrent tasks (concurrency)

paraformer-8k-v2

20

100

Fun-ASR

  • Singapore
  • China (Beijing)

Model name

Service deployment scope

Requests per minute (RPM)

fun-asr

International

600

fun-asr-2025-11-07

International

600

fun-asr-2025-08-25

International

600

fun-asr-mtl

International

100

fun-asr-mtl-2025-08-25

International

100

fun-asr-flash-2026-06-15

International

600

Fun-ASR-Realtime

  • Singapore
  • China (Beijing)

Model name

Service deployment scope

Job submission API RPS limit

fun-asr-realtime

International

20

fun-asr-realtime-2025-11-07

International

Text embedding

  • Singapore
  • China (Beijing)
  • Hong Kong (China)
Model nameService deployment scopeRate limiting conditions (triggered if any value is exceeded)
The following limits are per minute. The service may also enforce limits based on requests per second (RPS = RPM/60) and tokens per second (TPS = TPM/60).
Requests per minute (RPM)Tokens per minute (TPM)
Includes input and output tokens.
text-embedding-v4International1,8001,000,000
text-embedding-v3International6,00024,000,000

Multimodal embedding

  • Singapore
  • China (Beijing)
Model nameService deployment scopeRate limiting conditions
The following limits are per minute. The service may also enforce limits based on requests per second (RPS = RPM/60) and tokens per second (TPS = TPM/60).
Requests per minute (RPM)Tokens per minute (TPM)
Input tokens only.
tongyi-embedding-vision-plusInternational600200,000
tongyi-embedding-vision-flashInternational600200,000

Sorting model

  • Singapore
  • China (Beijing)
Model nameService deployment scopeRate limiting conditions
The following limits are per minute. The service may also enforce limits based on requests per second (RPS = RPM/60) and tokens per second (TPS = TPM/60).
Requests per minute (RPM)Tokens per minute (TPM)
Input tokens only.
qwen3-rerankInternational5,4005,000,000,000

Industry

Intention recognition

  • China (Beijing)
Model nameRate limiting conditions (triggered if any value is exceeded)
The following limits are per minute. The service may also enforce limits based on requests per second (RPS = RPM/60) and tokens per second (TPS = TPM/60).
Requests per minute (RPM)Tokens per minute (TPM)
Includes input and output tokens
tongyi-intent-detect-v31,2001,000,000

Role assumption

  • Singapore
  • China (Beijing)
Model nameService deployment scopeRate limiting conditions (triggered if any value is exceeded)
The following limits are per minute. The service may also enforce limits based on requests per second (RPS = RPM/60) and tokens per second (TPS = TPM/60).
Requests per minute (RPM)Tokens per minute (TPM)
Includes input and output tokens.
qwen-plus-characterInternational120500,000
qwen-flash-characterInternational120500,000
qwen-plus-character-jaInternational120500,000

Offline models

For more information, see Model unpublishing policy.
  • Offline on January 30, 2026
  • Offline on August 20, 2025
CategoryModel nameRate limiting conditions (triggered if any value is exceeded)
Requests per minute (RPM)Tokens per minute (TPM)
Includes input and output tokens.
Qwen-Plusqwen-plus-2024-11-2700
qwen-plus-2024-11-25
qwen-plus-2024-09-19
qwen-plus-2024-08-06
Qwen-Turboqwen-turbo-2024-09-19
Qwen-VLqwen-vl-max-2024-10-30
qwen-vl-max-2024-08-09
qwen-vl-plus-2024-08-09
Token Plan
Model Playground
  • Music generation
Statistics and Monitoring
Support