Skip to main content
Statistics and Monitoring

Model monitoring

Model Studio provides model monitoring, alerting, and logging so you can monitor model calls in real time, detect anomalies promptly, and troubleshoot issues.

Model monitoring overview

Model Studio provides unified model monitoring. A call is reflected in the monitoring charts about 1 to 2 minutes (up to 5 minutes) after it occurs, helping you track model calls in real time. Monitoring data is isolated by workspace: you can view data only for the currently selected workspace. Monitoring data is for reference only and is not used as a billing basis. For reconciliation, view your bill in Expenses and Costs.
Support for monitoring and alerting varies by model: some models—such as speech, image, and video generation models, and third-party models accessed directly—are not supported. Check the console to confirm.
This topic is organized by feature: to troubleshoot model calls, start with the audit log; to configure alerts, see the Alerting section; to enable data delivery, see the Data delivery section.

Filter monitoring data by dimension

You can filter monitoring data by Time (quick presets and custom), Inference Type, and API key (all by default; shows the ID and description, and lists only API keys that exist in the current workspace). Time-granularity filtering is available only on the model monitoring details page.

View overall call scale and health

The overview page shows data cards such as total models, total calls, total failures, average call duration, and average time to first token (first-token latency), so you can quickly gauge the overall call scale and health.

View monitoring details or logs for a model

The model list shows monitoring data for each model in the current workspace. Click View details to open the model monitoring details page, or click View Logs to open the Audit Log tab on the log page. Log on to the Model Studio console, and in the left-side navigation pane choose Model Monitoring & Alerting to open the monitoring overview page.

Model monitoring details

On the monitoring overview page, find the target model in the model list and click View details to open its model monitoring details page. The navigation path at the top of the page is: Model monitoring > model name. The details page shows call statistics and performance charts for a single model and provides an entry to configure alerts: the alert bell on a metric chart lets you quickly configure an alert for the current model or metric (for full cross-model rule management, see the Alerting section). The filters are the same as on the overview page (time, time granularity, inference type, and API key), and settings from the overview page are retained when you enter from it. To view logs, return to the model list on the overview page and click the View logs entry (see the Logging section).

Monitoring metrics

On the details page, metrics are shown in two groups—call statistics and performance. The meaning, reference notes, and alerting support of each metric are listed in the following table.

Metric

Description

Reference

Alerting support

Calls

Total number of times the model is called

Varies with business

Configurable

Usage

Total tokens

Varies with business

Configurable

Failures

Number of failed calls

Varies with business

Configurable

Failure rate

Failures divided by total calls

Varies with business

Configurable

Call duration

Total time from request to response

Varies with model and scenario

Configurable

Time to first token

Latency from the request to the first returned token

Varies with model and scenario

Configurable

Subsequent token latency

Average time to generate each token after the first

Varies with model and scenario

Not supported

Model TotalToken consumption

Total tokens consumed, aggregated by model (a metric dedicated to preset alert templates)

Varies with business

Configurable

Average usage per request

Average tokens per request

Varies with business

Not supported

Content moderation error count

Number of content moderation blocks

Varies with business

Not supported

Rate limiting error count

Number of 429 rate-limit triggers

Varies with business

Configurable

RPM

Requests per minute

Varies with business

Not supported

TPM

Tokens per minute

Varies with business

Not supported

Output tokens per second (TPS)

Output tokens per second

Varies with model and scenario

Not supported

Whether a metric supports alerting depends on whether it is included in a preset alert template (for preset templates, see Alerting > Alert templates). Metrics that are not in a preset template (such as content moderation error count and RPM) do not show an alert bell and cannot be configured for alerting.
For details about error codes and rate limiting, see Error codes.
Recommendation: To troubleshoot a surge in failed calls, focus on failures, rate limiting error count, and failure rate. To diagnose performance degradation, watch for a sustained upward trend in time to first token and average call duration, and configure alerts for key metrics.

Alert bell configuration

Metrics that support alerting show an alert bell icon on their chart. Click the bell to open the alert rule configuration side panel (for field details, see the Alert rule management section). The bell icon and the side-panel status labels have the following meanings:

Type

Status

Color

Meaning

Bell icon

No alert configured

Gray

No alert rule is configured; click to open the configuration side panel

Bell icon

Alert configured, normal

Blue

A rule is configured and not currently triggered; click to view rule details

Bell icon

Alerting

Red

An alert is being triggered; click to view trigger details

Side-panel label

Normal

Green

A rule is configured and its status is normal

Side-panel label

Alerting

Red

An alert is being triggered

Side-panel label

Disabled

Gray

The rule is turned off and no longer generates alerts

After you click the bell, the alert side panel shows different fields depending on the alert status (Normal = not triggered, Alerting = currently triggered, Recovered = triggered and then recovered):

Field

Normal

Recovered

Alerting

Alert rule name

Shown

Shown

Shown

Alert rule ID

Shown

Shown

Shown

Alert content

Shown

Shown

Shown

Alert metric

Shown

Shown

Shown

Notification recipients

Shown

Shown

Shown

Alert target

Shown

Shown

Shown

Start time

Hidden

Shown

Shown

Duration

Hidden

Shown

Shown

Alert count

Hidden

Shown

Shown

Status

Hidden

Shown

Shown

Recovery time

Hidden

Shown

Hidden

Alerting

Model Studio provides three alerting features—alert rules, alert templates, and alert history—all located under the Alert rules, Alert templates, and Alert history tabs on the Model Monitoring & Alerting page.

Alert rule management

On the Model Monitoring & Alerting page, click the Alerting tab to open the alert rule management page. The alert rule list shows the currently effective alert rules and supports creating, editing, stopping, and deleting rules. The list displays notification recipients (contacts or contact groups) and supports multiple channels such as email. The full path to set up alerting for the first time: first complete the CloudMonitor service-linked role authorization in Data delivery > Monitoring data delivery, then create an alert rule in this section (select an alert template and fill in the fields), and finally configure alert notifications (contacts or contact groups, notification time range, and repeat policy).
Recommendations: To control token costs, configure a spike alert (compared with the previous period) for Model TotalToken consumption. To ensure availability, configure an alert for the call failure rate. Both can directly use a preset alert template.
Before you create an alert rule, you must authorize the CloudMonitor service-linked role in model monitoring delivery (for how to configure it, see Data delivery > Monitoring data delivery). If it is not authorized, the Create alert rule button is not clickable, and hovering over it shows the tip: "The CloudMonitor service-linked role is not authorized, so you cannot create an alert rule. Click to authorize." After authorization, the CloudMonitor service is enabled automatically in the background.
  1. Click Create Alert Rule to open a new page.
  2. Fill in the alert name, alert template, model, duration, alert check period, alert content, and alert level as described in the field descriptions below.
  3. Configure alert notifications (alert contacts or contact groups, notification time range, and repeat policy), and then click Create to finish.
The form fields are described in the following table.

Field

Required

Description

Alert name

Required

Up to 50 characters.

Alert template

Required

Select a preset or custom template from the drop-down list, or create one by saving as based on a template.

Model

Required

Up to 50 models.

Duration

Required

Specified in minutes.

Alert check period

Required

Default 60 seconds, in seconds; must be an integer no less than 0 (0 means alert immediately when triggered).

Alert content

Required

Up to 200 characters; supports variables (you can insert placeholders such as workspace, model, and current value).

Alert level

Optional

INFO / WARNING / ERROR / CRITICAL. Default INFO.

Alert contacts / contact groups

Multiple allowed

Sourced from CloudMonitor contacts or contact groups.

Notification time range

Required

Any interval within 24 hours, and can span days (for example, 23:00 to 01:00 the next day).

Repeat policy

Required

No escalation means the alert is sent only once while unresolved; alternatively, specify an hour-plus-minute interval to repeat notifications.

For existing users who have already enabled monitoring delivery and configured alert rules, you can migrate alert rules from the CloudMonitor Prometheus instance to Model Studio with one click. Currently, only alert rules created by the platform can be migrated; alert rules created on the CloudMonitor side cannot be migrated.
  1. Above the alert rule list, click Migrate now or Migrate alert rules.
  2. Confirm the migration source. View the Prometheus instance name and status, and then click Start detection.
  3. Check alert rules. The system detects which alert rules can be migrated and which are incompatible. Select and confirm them, and then click Start migration.
  4. Migrate alert rules. During migration, a "Migrating" status is displayed. When migration finishes, the success or failure result is displayed. If migration fails, you can retry. After a successful migration, close the dialog box.
After migration, the alert rules in the Prometheus instance are disabled (not deleted); you can manually delete them on the Prometheus side. The migration task is retained for 6 hours, and refreshing the page during migration does not interrupt the task. After a successful migration, no second migration is needed, the migration entry is hidden, and only an inline success message is shown. Alert rules support the stop, edit, and delete operations. Alerting scope: currently only specific metrics support alerting. The configurable items are marked per row in the metric table in the Model monitoring details section.

Alert templates

An alert template is an alert configuration with preset trigger conditions. You can quickly create alert rules based on a template without configuring from scratch. The system provides 12 preset templates, and you can also create your own. The Alert templates tab is on the Model Monitoring & Alerting page, at the same level as the Alert rules tab and the Alert history tab. Click Create alert template to open the alert template creation side panel, and fill in the following information:
  • Template name: Required, up to 20 characters.
  • Parameter: Up to 10 parameters. Choose to trigger the alert when all rules meet the conditions (AND) or when any one meets the condition (OR).
  • Metric: Required. Selects the metric to monitor for alerting. The drop-down options include calls, failures, failure rate, call duration, and time to first token.
  • Alert Threshold: Required. Select a comparison operator (>, >=, <, <=, ==, !=) and a value.
  • Cycle: In minutes. Valid range 1 to 10080 minutes.
Preset alert templates are marked with a Preset tag after the name and support viewing details and copying, but not editing or deletion. Custom alert templates support viewing details, editing, deleting, and copying. Recommendations: use preset templates first to cover common alerting scenarios (the metric and threshold are preset for typical loads); just fill in the parameters to use them. Use custom templates for complex composite conditions and business-specific rules. The metrics supported by preset alert templates are listed in the following table.

Preset alert template name

Metric

Model call failures: 1-minute sum > 10

Model call failures

Model call failure ratio: 1-minute sum > 1%

Model call failure ratio

Model 4xx calls: 1-minute sum > 10

Model 4xx calls

Model 4xx ratio: 1-minute sum > 1%

Model 4xx ratio

Model 429 calls: 1-minute sum > 10

Model 429 calls

Model 429 ratio: 1-minute sum > 1%

Model 429 ratio

Model 5xx calls: 1-minute sum > 10

Model 5xx calls

Model 5xx ratio: 1-minute sum > 1%

Model 5xx ratio

Model calls: 1-minute sum > 500

Model calls

Model call duration: 1-minute average > 60 seconds

Model call duration

Model time to first token: 1-minute average > 10 seconds

Model time to first token (first-token latency)

Model TotalToken consumption: 1-minute sum > 10000

Model TotalToken consumption

Alert history

On the Model Monitoring & Alerting page, click the History tab to view alert trigger records. In the alert history list, the notification recipient shows only contact information, not the related notification channels. Alert history can be filtered by four conditions: alert time (quick presets and calendar), alert rule, alert level, and status (Recovered or Alerting). The alert history list displays the following columns.

Column

Description

Alert instance

Alert instance identifier

Alert level

INFO / WARNING / ERROR / CRITICAL

Alert time

When the alert was triggered

Alert count

Number of times this alert was triggered

Alert rule

Alert rule name and ID

Status

Recovered or Alerting

Notification recipient

Alert contacts or contact groups

Actions

View details (opens the alert details side panel)

Click View details in the list to open the alert details side panel. The status in the details side panel has only two values: Recovered and Alerting.

Logging

Model Studio records audit logs by default (stored on the platform, no configuration required), capturing the request information of each model call but not the Prompt or Response. Inference logs record the complete Prompt and Response and the intermediate steps; they must be enabled manually and rely on log delivery to write logs to your own Simple Log Service (SLS) Logstore (for log delivery, see the Data delivery section).

Audit Log

Audit logs are enabled by default for all users, require no configuration, and are stored on the Model Studio platform. Audit logs record the request information of each model call, including core metrics such as Request ID, time, model, token usage, latency, and status, but do not include the Prompt and Response content. To view the complete Prompt and Response content, enable inference logs (for how to enable them, see the Inference Log section).

Filter conditions

Audit logs support the following filter conditions:
  • Time range: Last 7 days by default; supports querying logs for up to 30 days.
  • Model: Drop-down selection; all models by default.
  • API key ID: Drop-down selection. API keys display the ID and description; if an API key is deleted, the description is empty.
  • Status: Multiple selection, including 0 (request succeeded but the user actively interrupted it), 200 (success), 4XX (exceptions caused by user behavior), and 5XX (service unavailable).
  • Inference Type: Real-time inference by default.
  • Search Request ID: Exact match.

Log list

The audit log list displays the following fields: Request ID, call time, model, usage, time to first token, call duration, status, and error code. The Actions column provides View details, which opens the audit log details side panel.

Details side panel

The audit log details side panel supports switching between the list view and the JSON view:
  • Basic information: Request ID, call time, model, API key.
  • Usage: total tokens, output tokens, input tokens.
  • Performance: time to first token, call duration, status code, error code, and error message (if any).
To create and manage API keys, go to System Management in the Model Studio console.

Inference Log

Inference logs must be enabled manually. On top of audit logs, inference logs additionally record the complete Prompt and Response and the intermediate steps, which help with debugging, optimization, and troubleshooting. Inference logs are stored in your own Simple Log Service Logstore.

Enable inference logs

Inference logs are disabled by default; you must first enable audit log delivery. On the log page, switch to the Inference Log tab (when it is not configured, this tab shows a guide page to enable it) and click Start configuration to open the log delivery configuration dialog box. You can also enter it by clicking the Log delivery configuration button in the upper-right corner of the log page. After you complete the authorization and Logstore configuration, you can view the inference log list.

Log list and details

The filter conditions for inference logs are the same as for audit logs, except that the inference type filter (real-time inference/batch inference) is removed. For models that do not support inference logs, "The current model is not supported" is displayed. On top of audit logs, the inference log list adds Prompt and Response data, with a length limit of 128 KB. Content beyond 128 KB is truncated; for the complete content, view the actual call logs.
Enabling inference logs and delivering logs to your own SLS Logstore incurs Simple Log Service charges. For the billing rules, see Billing overview of Simple Log Service.
The display rules after an API key is deleted are the same as for audit logs (the list does not show the description, and the details show only the ID). Inference logs can be recycled into training datasets. For more information, see Log backflow.
Recommendation: Inference logs record the complete Prompt and Response, which is suitable for reproducing abnormal requests, debugging model outputs, and troubleshooting intermittent issues.

Data delivery

By default, Model Studio stores monitoring metrics and audit logs on the platform, so you can view them in the console without any delivery. To deliver monitoring metrics or logs to your own Alibaba Cloud services (for long-term retention or integration with your own O&M systems, which incurs cloud service charges), enable the corresponding data delivery: monitoring data is delivered to a CloudMonitor Prometheus instance, and logs are delivered to a Simple Log Service (SLS) Logstore. Log delivery is also a prerequisite for inference logs.

Monitoring data delivery

The model monitoring delivery entry is in the upper-left corner of the monitoring overview page. Model monitoring delivery automatically delivers the monitoring data of Model Studio models to a Prometheus instance of Alibaba Cloud CloudMonitor for unified management of monitoring data. Model Studio provides model service metrics by default, which you can view on the platform without extra configuration. To deliver the metrics to the CloudMonitor service under your account, enable data delivery and configure the target resource.

Configuration steps

For new users, configuring model monitoring delivery requires the following steps:
  1. Authorize the CloudMonitor service-linked role.
  2. Activate the CloudMonitor service.
  3. Create a CloudMonitor Prometheus monitoring instance.
Model monitoring delivery currently supports only CloudMonitor Prometheus instances, not self-managed Prometheus. For existing users who have already enabled monitoring delivery, the existing configuration is migrated automatically, platform monitoring data is automatically delivered to the previously enabled Prometheus instance, and the Prometheus instance cannot be changed. After delivery is enabled, monitoring data is stored on both the platform side and the user side.
Existing users who have enabled advanced monitoring have monitoring data delivery enabled by default; if you do not need it, you can disable it at the model monitoring delivery entry.

Delivery status

After you configure data delivery, a status indicator is displayed next to model monitoring delivery:
  • Delivering (green)
  • Delivery failed (red)
  • Delivery not configured (gray)
The delivery toggle button is Call statistics and performance metrics delivery. Authorization details are expanded by default.

Log delivery

The log delivery configuration entry is in the upper-right corner of the log page. On the monitoring overview page, select a model to enter its details page, and then switch to the log page to see this entry. Log delivery automatically delivers the audit logs and inference logs of Model Studio models to a Logstore of Alibaba Cloud Simple Log Service for unified management of data. To deliver logs to the Simple Log Service under your account, enable log delivery and complete the configuration.

Configuration steps

Configuring log delivery for the first time requires the following three steps:
  1. Authorize the Simple Log Service role.
  2. Activate Simple Log Service.
  3. Create a Logstore.
Existing users who have already enabled log delivery do not need to reconfigure; the original configuration is retained automatically, and the Logstore cannot be changed.

Audit log delivery

After you enable or disable audit log delivery, a result message briefly pops up at the top of the page. The authorization and configuration process is the same as the configuration steps above.

Inference log delivery

Before you enable inference log delivery, you must first enable audit log delivery (inference log delivery cannot be enabled when audit log delivery is not configured). The authorization and configuration process for inference log delivery is the same as above.
After you disable audit log delivery, inference log delivery is also disabled.
After the configuration is complete, a status indicator is displayed next to the log delivery configuration; the three states have the same meanings as for monitoring data delivery.
After you disable log delivery, logs generated during the disabled period are not synchronized to the SLS Logstore and cannot be backfilled or restored afterward. Disable it with caution.
Token Plan
Model Playground
  • Music generation
Statistics and Monitoring
Support