GLM-5 dirancang untuk skenario coding dan agent, mencapai performa SOTA open-source dalam rekayasa sistem kompleks dengan kemampuan yang mendekati Claude Opus. Model ini dibangun di atas fondasi 744B menggunakan reinforcement learning asinkron dan sparse attention.
Inference Service Provider
Penyedia layanan inferensi untuk glm-5 adalah Alibaba Cloud Model Studio.
Model Capabilities
| Capability | Support | Capability | Support |
|---|---|---|---|
Input Modality | Text | Output Modality | Text |
Model Experience | Supported | Function Calling | Supported |
Structured Outputs | Unsupported | Web Search | Unsupported |
Prefix Completion | Unsupported | Context Caching | Supported |
Batch Inference | Unsupported | Fine-tuning | Unsupported |
Context Limits
| Parameter | Value | Parameter | Value |
|---|---|---|---|
Max Input Length | 169984 | Max Output Length | 16384 |
Context Window | 202752 | Max Input Length (Thinking Mode) | 169984 |
Max Output Length (Thinking Mode) | 16384 | Max Chain-of-Thought Length | 32768 |
Pricing
Halaman ini hanya menampilkan harga dasar untuk panggilan API model, tidak termasuk promosi berbatas waktu. Kunjungi Model Studio Console untuk penawaran promosi.
- China (Beijing)
Input<=32k
Input<=32k
| Billing Item | Price (USD) | Unit |
|---|---|---|
Input | 0,573 | Per 1M tokens |
Output | 2,58 | Per 1M tokens |
Input(Implicit Cache) | 0,115 | Per 1M tokens |
32k<Input<=200k
32k<Input<=200k
| Billing Item | Price (USD) | Unit |
|---|---|---|
Input | 0,86 | Per 1M tokens |
Output | 3,154 | Per 1M tokens |
Input(Implicit Cache) | 0,172 | Per 1M tokens |
Rate Limits
- China (Beijing)
| Parameter | Value |
|---|---|
RPM (Requests Per Minute) | 500 |
TPM (Tokens Per Minute) | 1.000.000 |