Model Studio provides model monitoring, alerting, and logging so you can monitor model calls in real time, detect anomalies promptly, and troubleshoot issues.
Model monitoring overview
Model Studio provides unified model monitoring. A call is reflected in the monitoring charts about 1 to 2 minutes (up to 5 minutes) after it occurs, helping you track model calls in real time. Monitoring data is isolated by workspace: you can view data only for the currently selected workspace. Monitoring data is for reference only and is not used as a billing basis. For reconciliation, view your bill in Expenses and Costs.
Filter monitoring data by dimension
You can filter monitoring data by Time (quick presets and custom), Inference Type, and API key (all by default; shows the ID and description, and lists only API keys that exist in the current workspace). Time-granularity filtering is available only on the model monitoring details page.
View overall call scale and health
The overview page shows data cards such as total models, total calls, total failures, average call duration, and average time to first token (first-token latency), so you can quickly gauge the overall call scale and health.
View monitoring details or logs for a model
The model list shows monitoring data for each model in the current workspace. Click View details to open the model monitoring details page, or click View Logs to open the Audit Log tab on the log page. Log on to the Model Studio console, and in the left-side navigation pane choose Model Monitoring & Alerting to open the monitoring overview page.
Model monitoring details
On the monitoring overview page, find the target model in the model list and click View details to open its model monitoring details page. The navigation path at the top of the page is: Model monitoring > model name.
The details page shows call statistics and performance charts for a single model and provides an entry to configure alerts: the alert bell on a metric chart lets you quickly configure an alert for the current model or metric (for full cross-model rule management, see the Alerting section). The filters are the same as on the overview page (time, time granularity, inference type, and API key), and settings from the overview page are retained when you enter from it. To view logs, return to the model list on the overview page and click the View logs entry (see the Logging section).
Monitoring metrics
On the details page, metrics are shown in two groups—call statistics and performance. The meaning, reference notes, and alerting support of each metric are listed in the following table.
Metric | Description | Reference | Alerting support |
|---|---|---|---|
Calls | Total number of times the model is called | Varies with business | Configurable |
Usage | Total tokens | Varies with business | Configurable |
Failures | Number of failed calls | Varies with business | Configurable |
Failure rate | Failures divided by total calls | Varies with business | Configurable |
Call duration | Total time from request to response | Varies with model and scenario | Configurable |
Time to first token | Latency from the request to the first returned token | Varies with model and scenario | Configurable |
Subsequent token latency | Average time to generate each token after the first | Varies with model and scenario | Not supported |
Model TotalToken consumption | Total tokens consumed, aggregated by model (a metric dedicated to preset alert templates) | Varies with business | Configurable |
Average usage per request | Average tokens per request | Varies with business | Not supported |
Content moderation error count | Number of content moderation blocks | Varies with business | Not supported |
Rate limiting error count | Number of 429 rate-limit triggers | Varies with business | Configurable |
RPM | Requests per minute | Varies with business | Not supported |
TPM | Tokens per minute | Varies with business | Not supported |
Output tokens per second (TPS) | Output tokens per second | Varies with model and scenario | Not supported |
Alert bell configuration
Metrics that support alerting show an alert bell icon on their chart. Click the bell to open the alert rule configuration side panel (for field details, see the Alert rule management section). The bell icon and the side-panel status labels have the following meanings:
Type | Status | Color | Meaning |
|---|---|---|---|
Bell icon | No alert configured | Gray | No alert rule is configured; click to open the configuration side panel |
Bell icon | Alert configured, normal | Blue | A rule is configured and not currently triggered; click to view rule details |
Bell icon | Alerting | Red | An alert is being triggered; click to view trigger details |
Side-panel label | Normal | Green | A rule is configured and its status is normal |
Side-panel label | Alerting | Red | An alert is being triggered |
Side-panel label | Disabled | Gray | The rule is turned off and no longer generates alerts |
Field | Normal | Recovered | Alerting |
|---|---|---|---|
Alert rule name | Shown | Shown | Shown |
Alert rule ID | Shown | Shown | Shown |
Alert content | Shown | Shown | Shown |
Alert metric | Shown | Shown | Shown |
Notification recipients | Shown | Shown | Shown |
Alert target | Shown | Shown | Shown |
Start time | Hidden | Shown | Shown |
Duration | Hidden | Shown | Shown |
Alert count | Hidden | Shown | Shown |
Status | Hidden | Shown | Shown |
Recovery time | Hidden | Shown | Hidden |
Alerting
Model Studio provides three alerting features—alert rules, alert templates, and alert history—all located under the Alert rules, Alert templates, and Alert history tabs on the Model Monitoring & Alerting page.
Alert rule management
On the Model Monitoring & Alerting page, click the Alerting tab to open the alert rule management page. The alert rule list shows the currently effective alert rules and supports creating, editing, stopping, and deleting rules. The list displays notification recipients (contacts or contact groups) and supports multiple channels such as email.
The full path to set up alerting for the first time: first complete the CloudMonitor service-linked role authorization in Data delivery > Monitoring data delivery, then create an alert rule in this section (select an alert template and fill in the fields), and finally configure alert notifications (contacts or contact groups, notification time range, and repeat policy).
- Click Create Alert Rule to open a new page.
- Fill in the alert name, alert template, model, duration, alert check period, alert content, and alert level as described in the field descriptions below.
- Configure alert notifications (alert contacts or contact groups, notification time range, and repeat policy), and then click Create to finish.
Field | Required | Description |
|---|---|---|
Alert name | Required | Up to 50 characters. |
Alert template | Required | Select a preset or custom template from the drop-down list, or create one by saving as based on a template. |
Model | Required | Up to 50 models. |
Duration | Required | Specified in minutes. |
Alert check period | Required | Default 60 seconds, in seconds; must be an integer no less than 0 (0 means alert immediately when triggered). |
Alert content | Required | Up to 200 characters; supports variables (you can insert placeholders such as workspace, model, and current value). |
Alert level | Optional | INFO / WARNING / ERROR / CRITICAL. Default INFO. |
Alert contacts / contact groups | Multiple allowed | Sourced from CloudMonitor contacts or contact groups. |
Notification time range | Required | Any interval within 24 hours, and can span days (for example, 23:00 to 01:00 the next day). |
Repeat policy | Required | No escalation means the alert is sent only once while unresolved; alternatively, specify an hour-plus-minute interval to repeat notifications. |
- Above the alert rule list, click Migrate now or Migrate alert rules.
- Confirm the migration source. View the Prometheus instance name and status, and then click Start detection.
- Check alert rules. The system detects which alert rules can be migrated and which are incompatible. Select and confirm them, and then click Start migration.
- Migrate alert rules. During migration, a "Migrating" status is displayed. When migration finishes, the success or failure result is displayed. If migration fails, you can retry. After a successful migration, close the dialog box.
Alert templates
An alert template is an alert configuration with preset trigger conditions. You can quickly create alert rules based on a template without configuring from scratch. The system provides 12 preset templates, and you can also create your own. The Alert templates tab is on the Model Monitoring & Alerting page, at the same level as the Alert rules tab and the Alert history tab.
Click Create alert template to open the alert template creation side panel, and fill in the following information:
- Template name: Required, up to 20 characters.
- Parameter: Up to 10 parameters. Choose to trigger the alert when all rules meet the conditions (AND) or when any one meets the condition (OR).
- Metric: Required. Selects the metric to monitor for alerting. The drop-down options include calls, failures, failure rate, call duration, and time to first token.
- Alert Threshold: Required. Select a comparison operator (>, >=, <, <=, ==, !=) and a value.
- Cycle: In minutes. Valid range 1 to 10080 minutes.
Preset alert template name | Metric |
|---|---|
Model call failures: 1-minute sum > 10 | Model call failures |
Model call failure ratio: 1-minute sum > 1% | Model call failure ratio |
Model 4xx calls: 1-minute sum > 10 | Model 4xx calls |
Model 4xx ratio: 1-minute sum > 1% | Model 4xx ratio |
Model 429 calls: 1-minute sum > 10 | Model 429 calls |
Model 429 ratio: 1-minute sum > 1% | Model 429 ratio |
Model 5xx calls: 1-minute sum > 10 | Model 5xx calls |
Model 5xx ratio: 1-minute sum > 1% | Model 5xx ratio |
Model calls: 1-minute sum > 500 | Model calls |
Model call duration: 1-minute average > 60 seconds | Model call duration |
Model time to first token: 1-minute average > 10 seconds | Model time to first token (first-token latency) |
Model TotalToken consumption: 1-minute sum > 10000 | Model TotalToken consumption |
Alert history
On the Model Monitoring & Alerting page, click the History tab to view alert trigger records. In the alert history list, the notification recipient shows only contact information, not the related notification channels.
Alert history can be filtered by four conditions: alert time (quick presets and calendar), alert rule, alert level, and status (Recovered or Alerting).
The alert history list displays the following columns.
Column | Description |
|---|---|
Alert instance | Alert instance identifier |
Alert level | INFO / WARNING / ERROR / CRITICAL |
Alert time | When the alert was triggered |
Alert count | Number of times this alert was triggered |
Alert rule | Alert rule name and ID |
Status | Recovered or Alerting |
Notification recipient | Alert contacts or contact groups |
Actions | View details (opens the alert details side panel) |
Logging
Model Studio records audit logs by default (stored on the platform, no configuration required), capturing the request information of each model call but not the Prompt or Response. Inference logs record the complete Prompt and Response and the intermediate steps; they must be enabled manually and rely on log delivery to write logs to your own Simple Log Service (SLS) Logstore (for log delivery, see the Data delivery section).
Audit Log
Audit logs are enabled by default for all users, require no configuration, and are stored on the Model Studio platform. Audit logs record the request information of each model call, including core metrics such as Request ID, time, model, token usage, latency, and status, but do not include the Prompt and Response content. To view the complete Prompt and Response content, enable inference logs (for how to enable them, see the Inference Log section).
Filter conditions
Audit logs support the following filter conditions:
- Time range: Last 7 days by default; supports querying logs for up to 30 days.
- Model: Drop-down selection; all models by default.
- API key ID: Drop-down selection. API keys display the ID and description; if an API key is deleted, the description is empty.
- Status: Multiple selection, including 0 (request succeeded but the user actively interrupted it), 200 (success), 4XX (exceptions caused by user behavior), and 5XX (service unavailable).
- Inference Type: Real-time inference by default.
- Search Request ID: Exact match.
Log list
The audit log list displays the following fields: Request ID, call time, model, usage, time to first token, call duration, status, and error code. The Actions column provides View details, which opens the audit log details side panel.
Details side panel
The audit log details side panel supports switching between the list view and the JSON view:
- Basic information: Request ID, call time, model, API key.
- Usage: total tokens, output tokens, input tokens.
- Performance: time to first token, call duration, status code, error code, and error message (if any).
Inference Log
Inference logs must be enabled manually. On top of audit logs, inference logs additionally record the complete Prompt and Response and the intermediate steps, which help with debugging, optimization, and troubleshooting. Inference logs are stored in your own Simple Log Service Logstore.
Enable inference logs
Inference logs are disabled by default; you must first enable audit log delivery. On the log page, switch to the Inference Log tab (when it is not configured, this tab shows a guide page to enable it) and click Start configuration to open the log delivery configuration dialog box. You can also enter it by clicking the Log delivery configuration button in the upper-right corner of the log page. After you complete the authorization and Logstore configuration, you can view the inference log list.
Log list and details
The filter conditions for inference logs are the same as for audit logs, except that the inference type filter (real-time inference/batch inference) is removed. For models that do not support inference logs, "The current model is not supported" is displayed.
On top of audit logs, the inference log list adds Prompt and Response data, with a length limit of 128 KB. Content beyond 128 KB is truncated; for the complete content, view the actual call logs.
Data delivery
By default, Model Studio stores monitoring metrics and audit logs on the platform, so you can view them in the console without any delivery. To deliver monitoring metrics or logs to your own Alibaba Cloud services (for long-term retention or integration with your own O&M systems, which incurs cloud service charges), enable the corresponding data delivery: monitoring data is delivered to a CloudMonitor Prometheus instance, and logs are delivered to a Simple Log Service (SLS) Logstore. Log delivery is also a prerequisite for inference logs.
Monitoring data delivery
The model monitoring delivery entry is in the upper-left corner of the monitoring overview page. Model monitoring delivery automatically delivers the monitoring data of Model Studio models to a Prometheus instance of Alibaba Cloud CloudMonitor for unified management of monitoring data. Model Studio provides model service metrics by default, which you can view on the platform without extra configuration. To deliver the metrics to the CloudMonitor service under your account, enable data delivery and configure the target resource.
Configuration steps
For new users, configuring model monitoring delivery requires the following steps:
- Authorize the CloudMonitor service-linked role.
- Activate the CloudMonitor service.
- Create a CloudMonitor Prometheus monitoring instance.
Delivery status
After you configure data delivery, a status indicator is displayed next to model monitoring delivery:
- Delivering (green)
- Delivery failed (red)
- Delivery not configured (gray)
Log delivery
The log delivery configuration entry is in the upper-right corner of the log page. On the monitoring overview page, select a model to enter its details page, and then switch to the log page to see this entry. Log delivery automatically delivers the audit logs and inference logs of Model Studio models to a Logstore of Alibaba Cloud Simple Log Service for unified management of data. To deliver logs to the Simple Log Service under your account, enable log delivery and complete the configuration.
Configuration steps
Configuring log delivery for the first time requires the following three steps:
- Authorize the Simple Log Service role.
- Activate Simple Log Service.
- Create a Logstore.