Qwen3.8-2.4T-A95B is the open-source release of Qwen's latest flagship, launched August 2026. Its sparse MoE architecture holds 2.4T total parameters with ~95B activated per step, paired with hybrid attention and a 1M token context window. Key benchmarks: GPQA Diamond 92.6, PaperBench 93.0, OSWorld 86.1, BabyVision 82.0. Ranked 4th on CodeArena.
Inference Service Provider
The inference service provider for qwen3.8-2.4t-a95b is Alibaba Cloud Model Studio.
Model Capabilities
| Capability | Support | Capability | Support |
|---|---|---|---|
Input Modality | Text | Output Modality | Text |
Model Experience | Supported | Function Calling | Supported |
Structured Outputs | Supported | Web Search | Supported |
Prefix Completion | Supported | Context Caching | Supported |
Batch Inference | Supported | Fine-tuning | Unsupported |
Context Limits
Parameter | Value | Parameter | Value |
|---|---|---|---|
Max Input Length | 991808 | Max Output Length | 131072 |
Context Length | 1000000 | Max Input Length (Thinking Mode) | 983616 |
Max Output Length (Thinking Mode) | 131072 | Max Chain-of-Thought Length | 131072 |
Pricing
This page only shows the original pricing for model API calls, excluding any limited-time promotions. Visit Model Studio Console for promotional offers.
- China (Beijing)
- Singapore
| Billing Item | Price (USD) | Unit |
|---|---|---|
Input | 1.65 | Per 1M tokens |
Output | 4.951 | Per 1M tokens |
Input (Cache Hit) | 0.206 | Per 1M tokens |
Explicit Cache Creation | 2.063 | Per 1M tokens |
Explicit Cache Hit | 0.137 | Per 1M tokens |
Rate Limits
- China (Beijing)
- Singapore
| Parameter | Value |
|---|---|
RPM (Requests Per Minute) | 5000 |
TPM (Tokens Per Minute) | 5,000,000 |