Recommended Metrics and Alarm Policies for Each Cloud Service
This section recommends metrics and alarm policies for configuring alarms for specific cloud services. Alarm policies change based on cloud services. This section is for reference only. Adjust your alarm settings as required.
Elastic Cloud Server
| Namespace | Metric Name | Metric ID | Statistic | Consecutive Triggers | Operator | Major Alarm Threshold | Critical Alarm Threshold | Unit | Frequency |
|---|---|---|---|---|---|---|---|---|---|
| SYS.ECS | CPU Usage | cpu_util | Raw data | 3 | > | 80 | 90 | % | Hourly |
| (Windows) Memory Usage | mem_util | Raw data | 3 | > | 80 | 90 | % | Hourly | |
| (Windows) Disk Usage | disk_util_inband | Raw data | 3 | > | 80 | 90 | % | Hourly | |
| AGT.ECS | (Agent) CPU Usage | cpu_usage | Raw data | 3 | > | 80 | 90 | % | Hourly |
| (Agent) Memory Usage | mem_usedPercent | Raw data | 1 | > | 80 | 90 | % | Hourly | |
| (Agent) Receive Error Rate | net_errin | Raw data | 5 | > | 0 | - | % | Every 5 minutes | |
| (Agent) Transmit Error Rate | net_errout | Raw data | 5 | > | 0 | - | % | Every 5 minutes | |
| (Agent) Received Packet Drop Rate | net_dropin | Raw data | 5 | > | 0 | - | % | Every 5 minutes | |
| (Agent) Transmitted Packet Drop Rate | net_dropout | Raw data | 5 | > | 0 | - | % | Every 5 minutes | |
| (Agent) Blocked Processes | proc_blocked_count | Raw data | 5 | > | 0 | - | count | Hourly | |
| (Agent) NTP Offset | ntp_offset | Raw data | 3 | >= | 5000 | 10000 | ms | Hourly | |
| (Agent) Disk I/O Usage | disk_ioUtils | Raw data | 3 | > | 80 | 90 | % | Hourly | |
| (Agent) Disk Usage | disk_usedPercent | Raw data | 3 | > | 80 | 90 | % | Hourly | |
| (Agent) Percentage of Total inode Used | disk_inodesUsedPercent | Raw data | 3 | > | 80 | 90 | % | Hourly | |
| (Agent) File System Read/Write Status | disk_fs_rwstate | Raw data | 2 | = | - | 1 | N/A | Hourly | |
| (Agent) NPU Device Health | npu_device_health | Raw data | 1 | = | 2 | 3 | N/A | Hourly | |
| (Agent) NPU Driver Health | npu_driver_health | Raw data | 5 | != | - | 0 | N/A | Once | |
| (Agent) NPU Memory Usage | npu_util_rate_mem | Raw data | 5 | > | 98 | - | % | Once | |
| (Agent) NPU AI Core Usage | npu_util_rate_ai_core | Raw data | 10 | > | 98 | - | % | Once | |
| (Agent) NPU Control CPU Usage | npu_util_rate_ctrl_cpu | Raw data | 10 | > | 98 | - | % | Once | |
| (Agent) Average NPU AI CPU Usage | npu_aicpu_avg_util_rate | Raw data | 10 | > | 98 | - | % | Once | |
| (Agent) HBM ECC Check Status | npu_hbm_ecc_enable | Raw data | 5 | = | 0 | - | N/A | Once | |
| (Agent) Isolated Memory Pages with HBM Double-Bit Errors | npu_hbm_double_bit_isolated_pages_cnt | Raw data | 5 | >= | 64 | - | count | Once | |
| (Agent) NPU HBM Usage | npu_util_rate_hbm | Raw data | 5 | > | 95 | 98 | % | Once | |
| (Agent) NPU Optical Module Case Temperature | npu_opt_temperature | Raw data | 5 | > < | - | 80 -10 | °C | Once | |
| NPU Vector Core Usage | npu_util_rate_vector_core | Raw data | 10 | > | 98 | - | % | Once | |
| NPU Macro1 SerDes Lane 0 SNR | npu_macro1_serdes_lane0_snr | Raw data | 5 | < | - | 500000 | db | Once | |
| NPU Macro1 SerDes Lane 1 SNR | npu_macro1_serdes_lane1_snr | Raw data | 5 | < | - | 500000 | db | Once | |
| NPU Macro1 SerDes Lane 2 SNR | npu_macro1_serdes_lane2_snr | Raw data | 5 | < | - | 500000 | db | Once | |
| NPU Macro1 SerDes Lane 3 SNR | npu_macro1_serdes_lane3_snr | Raw data | 5 | < | - | 500000 | db | Once | |
| NPU Macro2 SerDes Lane 0 SNR | npu_macro2_serdes_lane0_snr | Raw data | 5 | < | - | 500000 | db | Once | |
| NPU Macro2 SerDes Lane 1 SNR | npu_macro2_serdes_lane1_snr | Raw data | 5 | < | - | 500000 | db | Once | |
| NPU Macro2 SerDes Lane 2 SNR | npu_macro2_serdes_lane2_snr | Raw data | 5 | < | - | 500000 | db | Once | |
| NPU Macro2 SerDes Lane 3 SNR | npu_macro2_serdes_lane3_snr | Raw data | 5 | < | - | 500000 | db | Once | |
| NPU Macro3 SerDes Lane 0 SNR | npu_macro3_serdes_lane0_snr | Raw data | 5 | < | - | 500000 | db | Once | |
| NPU Macro3 SerDes Lane 1 SNR | npu_macro3_serdes_lane1_snr | Raw data | 5 | < | - | 500000 | db | Once | |
| NPU Macro3 SerDes Lane 2 SNR | npu_macro3_serdes_lane2_snr | Raw data | 5 | < | - | 500000 | db | Once | |
| NPU Macro3 SerDes Lane 3 SNR | npu_macro3_serdes_lane3_snr | Raw data | 5 | < | - | 500000 | db | Once | |
| NPU Macro4 SerDes Lane 0 SNR | npu_macro4_serdes_lane0_snr | Raw data | 5 | < | - | 500000 | db | Once | |
| NPU Macro4 SerDes Lane 1 SNR | npu_macro4_serdes_lane1_snr | Raw data | 5 | < | - | 500000 | db | Once | |
| NPU Macro4 SerDes Lane 2 SNR | npu_macro4_serdes_lane2_snr | Raw data | 5 | < | - | 500000 | db | Once | |
| NPU Macro4 SerDes Lane 3 SNR | npu_macro4_serdes_lane3_snr | Raw data | 5 | < | - | 500000 | db | Once | |
| NPU Macro5 SerDes Lane 0 SNR | npu_macro5_serdes_lane0_snr | Raw data | 5 | < | - | 500000 | db | Once | |
| NPU Macro5 SerDes Lane 1 SNR | npu_macro5_serdes_lane1_snr | Raw data | 5 | < | - | 500000 | db | Once | |
| NPU Macro5 SerDes Lane 2 SNR | npu_macro5_serdes_lane2_snr | Raw data | 5 | < | - | 500000 | db | Once | |
| NPU Macro5 SerDes Lane 3 SNR | npu_macro5_serdes_lane3_snr | Raw data | 5 | < | - | 500000 | db | Once | |
| NPU Macro6 SerDes Lane 0 SNR | npu_macro6_serdes_lane0_snr | Raw data | 5 | < | - | 500000 | db | Once | |
| NPU Macro6 SerDes Lane 1 SNR | npu_macro6_serdes_lane1_snr | Raw data | 5 | < | - | 500000 | db | Once | |
| NPU Macro6 SerDes Lane 2 SNR | npu_macro6_serdes_lane2_snr | Raw data | 5 | < | - | 500000 | db | Once | |
| NPU Macro6 SerDes Lane 3 SNR | npu_macro6_serdes_lane3_snr | Raw data | 5 | < | - | 500000 | db | Once | |
| NPU Macro7 SerDes Lane 0 SNR | npu_macro7_serdes_lane0_snr | Raw data | 5 | < | - | 500000 | db | Once | |
| NPU Macro7 SerDes Lane 1 SNR | npu_macro7_serdes_lane1_snr | Raw data | 5 | < | - | 500000 | db | Once | |
| NPU Macro7 SerDes Lane 2 SNR | npu_macro7_serdes_lane2_snr | Raw data | 5 | < | - | 500000 | db | Once | |
| NPU Macro7 SerDes Lane 3 SNR | npu_macro7_serdes_lane3_snr | Raw data | 5 | < | - | 500000 | db | Once | |
| Packets Retransmitted by NPU RoCE | npu_roce_new_pkt_rty_num | Raw data | 5 | Increase compared with last period | 1 | - | % | Once | |
| Abnormal PSN Packets Received by NPU RoCE | npu_roce_out_of_order_num | Raw data | 5 | Increase compared with last period | 1 | - | % | Once |
API Gateway (Dedicated)
| Namespace | Metric Name | Metric ID | Statistic | Consecutive Triggers | Operator | Major Alarm Threshold | Critical Alarm Threshold | Unit | Frequency |
|---|---|---|---|---|---|---|---|---|---|
| SYS.APIC | 5xx Responses | req_count_5xx | Raw data | 1 | Increase compared with last period | 20 | 30 | % | Hourly |
| Average Latency | avg_latency | Raw data | 3 | >= | 3000 | 5000 | ms | Hourly | |
| Node System Load | node_system_load | Raw data | 3 | = | 2 | 3 | count | Hourly | |
| Node CPU Usage | node_cpu_usage | Raw data | 3 | > | 30 | 60 | % | Hourly | |
| Node Memory Usage | node_memory_usage | Raw data | 3 | > | 30 | 60 | % | Hourly | |
| Throttled API Calls | throttled_calls | Raw data | 1 | Increase compared with last period | 50 | 70 | % | Hourly |
NAT Gateway
| Namespace | Metric Name | Metric ID | Statistic | Consecutive Triggers | Operator | Major Alarm Threshold | Critical Alarm Threshold | Unit | Frequency |
|---|---|---|---|---|---|---|---|---|---|
| SYS.NAT | Inbound PPS | inbound_pps | Raw data | 3 | > | - | 800000 | Count | Hourly |
| Inbound PPS | inbound_pps | Raw data | 3 | Increase or decrease compared with last period | 20 | - | % | Hourly | |
| Outbound PPS | outbound_pps | Raw data | 3 | > | - | > 800000 | Count | Hourly | |
| Outbound PPS | outbound_pps | Raw data | 3 | Increase or decrease compared with last period | 20 | - | % | Hourly | |
| SNAT Connection Usage Rate | snat_connection_ratio | Raw data | 3 | > | - | 80 | % | Hourly | |
| Packets Dropped (Excessive SNAT Connections) | packets_drop_count_snat_connection_beyond | Raw data | 3 | > | - | 0 | Count | Hourly | |
| Packets Dropped (Excessive PPS) | packets_drop_count_pps_beyond | Raw data | 3 | > | - | 0 | Count | Hourly | |
| Packets Dropped (When All EIP Ports Allocated) | packets_drop_count_eip_port_alloc_beyond | Raw data | 3 | > | - | 0 | Count | Hourly | |
| Total PPS Usage | total_pps_ratio | Raw data | 3 | > | - | 80 | % | Hourly |
Web Application Firewall
| Namespace | Metric Name | Metric ID | Statistic | Consecutive Triggers | Operator | Major Alarm Threshold | Critical Alarm Threshold | Unit | Frequency |
|---|---|---|---|---|---|---|---|---|---|
| SYS.WAF | CPU Usage | cpu_util | Raw data | 3 | > | 80 | 90 | % | Hourly |
| Memory Usage | mem_util | Raw data | 3 | > | 80 | 90 | % | Hourly | |
| Disk Usage | disk_util | Raw data | 3 | > | 80 | - | % | Hourly | |
| Active Connections | active_connections | Raw data | 3 | > | 40000 | - | Count | Hourly | |
| WAF Status Code (5XX) | waf_http_5xx | Raw data | 1 | Increase compared with last period | 10 | 15 | % | Hourly | |
| Status Code Returned by the Origin Server (5XX) | upstream_code_5xx | Raw data | 3 | > | 15 | 20 | Count | Hourly |
Elastic Load Balance
| Namespace | Metric Name | Metric ID | Statistic | Consecutive Triggers | Operator | Major Alarm Threshold | Critical Alarm Threshold | Unit | Frequency |
|---|---|---|---|---|---|---|---|---|---|
| SYS.ELB | Concurrent Connections | m1_cps | Raw data | 3 | > | 40000 | 45000 | Count | Hourly |
| New Connections | m4_ncps | Raw data | 3 | > | 4000 | 4500 | Count/s | Hourly | |
| Unhealthy Servers | m9_abnormal_servers | Raw data | 3 | > | - | 0 | Count | Hourly | |
| Dropped Connections | dropped_connections | Raw data | 3 | > | - | 0 | Count/s | Hourly | |
| Dropped Packets | dropped_packets | Raw data | 3 | > | - | 0 | Count/s | Hourly | |
| Bandwidth for Dropping Packets | dropped_traffic | Raw data | 3 | > | - | 0 | bit/s | Hourly | |
| Layer 4 New Connection Usage | l4_ncps_usage | Raw data | 3 | > | 80 | 90 | % | Hourly | |
| Layer 4 Concurrent Connection Usage | l4_con_usage | Raw data | 3 | > | 80 | 90 | % | Hourly | |
| Layer 4 Inbound Bandwidth Usage | l4_in_bps_usage | Raw data | 3 | > | 80 | 90 | % | Hourly | |
| Layer 4 Outbound Bandwidth Usage | l4_out_bps_usage | Raw data | 3 | > | 80 | 90 | % | Hourly | |
| Layer 7 New Connection Usage | l7_ncps_usage | Raw data | 3 | > | 80 | 90 | % | Hourly | |
| Layer 7 Concurrent Connection Usage | l7_con_usage | Raw data | 3 | > | 80 | 90 | % | Hourly | |
| Layer 7 Inbound Bandwidth Usage | l7_in_bps_usage | Raw data | 3 | > | 80 | 90 | % | Hourly | |
| Layer 7 Outbound Bandwidth Usage | l7_out_bps_usage | Raw data | 3 | > | 80 | 90 | % | Hourly | |
| Layer 7 QPS Usage | l7_qps_usage | Raw data | 3 | > | 80 | 90 | % | Hourly | |
| Concurrent Connections | m1_cps | Raw data | 1 | Decrease compared with previous period | - | 80 | % | Hourly | |
| New Connections | m4_ncps | Raw data | 1 | Decrease compared with previous period | - | 80 | % | Hourly | |
| 5xx Status Codes (Total) | mf_l7_http_5xx | Raw data | 1 | Increase compared with last period | - | 50 | % | Hourly | |
| Average Layer 7 Response Time | m14_l7_rt | Raw data | 1 | Increase compared with last period | - | 50 | % | Hourly | |
| 5xx Status Codes (Load Balancer) | elb_http_5xx | Raw data | 1 | Increase compared with last period | - | 50 | % | Hourly | |
| 5xx Status Codes Percentage | l7_5xx_ratio | Raw data | 3 | >= | - | 5 | % | Hourly | |
| 2xx Status Codes Percentage | l7_2xx_ratio | Raw data | 3 | <= | - | 95 | % | Hourly |
Scalable File Service
| Namespace | Metric Name | Metric ID | Statistic | Consecutive Triggers | Operator | Major Alarm Threshold | Critical Alarm Threshold | Unit | Frequency |
|---|---|---|---|---|---|---|---|---|---|
| SYS.SFS | File System Read Bandwidth | read_bytes_intranet | Raw data | 1 | Decrease compared with previous period | 100 | - | % | Every 3 hours |
| File System Write Bandwidth | write_bytes_intranet | Raw data | 1 | Decrease compared with previous period | 100 | - | % | Every 3 hours | |
| File System Read TPS | read_tps | Raw data | 1 | Decrease compared with previous period | 100 | - | % | Every 3 hours | |
| File System Write TPS | write_tps | Raw data | 1 | Decrease compared with previous period | 100 | - | % | Every 3 hours |
Scalable File Service Turbo
| Namespace | Metric Name | Metric ID | Statistic | Consecutive Triggers | Operator | Major Alarm Threshold | Critical Alarm Threshold | Unit | Frequency |
|---|---|---|---|---|---|---|---|---|---|
| SYS.EFS | Capacity Usage | used_capacity_percent | Raw data | 5 | > | 90 | 95 | % | Hourly |
| Inode Usage | used_inode_percent | Raw data | 5 | > | 90 | 95 | % | Hourly |
Object Storage Service
| Namespace | Metric Name | Metric ID | Statistic | Consecutive Triggers | Operator | Major Alarm Threshold | Critical Alarm Threshold | Unit | Frequency |
|---|---|---|---|---|---|---|---|---|---|
| SYS.OBS | Request Success Rate | request_success_rate | Raw data | 2 | < | - | 99.97 | % | Hourly |
Distributed Cache Service
| Namespace | Metric Name | Metric ID | Statistic | Consecutive Triggers | Operator | Major Alarm Threshold | Critical Alarm Threshold | Unit | Frequency |
|---|---|---|---|---|---|---|---|---|---|
| SYS.DCS | Memory Usage | memory_usage | Raw data | 2 | > | 70 | 80 | % | Hourly |
| CPU Usage | cpu_usage | Raw data | 2 | Decrease compared with previous period | - | 100 | % | Hourly | |
| Proxy Status | node_status | Raw data | 2 | = | - | 1 | N/A | Hourly | |
| Average CPU Usage | cpu_avg_usage | Raw data | 2 | > | 70 | 80 | % | Hourly | |
| Maximum Latency | command_max_rt | Raw data | 2 | > | - | 900000 | μs | Hourly | |
| Average Latency | command_avg_rt | Raw data | 2 | > | - | 150000 | μs | Hourly | |
| Connection Usage | connections_usage | Raw data | 2 | > | 70 | 80 | % | Hourly | |
| CPU Usage | cpu_usage | Raw data | 2 | > | 70 | 80 | % | Hourly | |
| Slow Query Logs | mc_is_slow_log_exist | Raw data | 1 | > | - | 0 | N/A | Hourly |
Distributed Database Middleware
| Namespace | Metric Name | Metric ID | Statistic | Consecutive Triggers | Operator | Major Alarm Threshold | Critical Alarm Threshold | Unit | Frequency |
|---|---|---|---|---|---|---|---|---|---|
| SYS.DDMS | CPU Usage | ddm_cpu_util | Raw data | 3 | > | 80 | 90 | % | Hourly |
| Memory Usage | ddm_mem_util | Raw data | 3 | > | 85 | 90 | % | Hourly | |
| Slow SQL Logs | ddm_slow_log | Raw data | 3 | > | 50 | 100 | Piece | Daily | |
| Connection Usage | ddm_connection_util | Raw data | 2 | >= | 80 | 85 | % | Hourly | |
| DDM Node Connectivity | ddm_node_status_alarm_code | Raw data | 1 | = | - | 1 | N/A | Hourly |
Distributed Message Service
| Namespace | Metric Name | Metric ID | Statistic | Consecutive Triggers | Operator | Major Alarm Threshold | Critical Alarm Threshold | Unit | Frequency |
|---|---|---|---|---|---|---|---|---|---|
| SYS.DMS | Consumers | consumers | Raw data | 2 | > | 3600 | - | Count | Hourly |
| RabbitMQ Instance Available Messages | messages_ready | Raw data | 1 | > | 10000 | - | Count | Hourly | |
| Unacknowledged Messages | messages_unacknowledged | Raw data | 1 | > | 10000 | - | Count | Hourly | |
| Instance Disk Usage | instance_disk_usage | Raw data | 3 | > | 80 | 90 | % | Hourly | |
| Disk Capacity Usage | broker_disk_usage | Raw data | 3 | > | 80 | 90 | % | Hourly | |
| Memory Usage | broker_memory_usage | Raw data | 3 | > | 80 | 90 | % | Hourly | |
| Broker Alive | broker_alive | Raw data | 1 | = | - | 0 | N/A | Hourly | |
| Connections | broker_connections | Raw data | 3 | > | - | 2000 | Count | Hourly | |
| CPU Usage | broker_cpu_usage | Raw data | 3 | Decrease compared with previous period | - | 100 | % | Hourly | |
| Average Disk Read Time | broker_disk_read_await | Raw data | 3 | > | - | 5000 | ms | Hourly | |
| Average Disk Write Time | broker_disk_write_await | Raw data | 3 | > | - | 5000 | ms | Hourly | |
| Message Creation Processing (99th Percentile) | broker_produce_p99 | Raw data | 3 | > | 50 | - | ms | Hourly | |
| Message Creation Processing (99.9th Percentile) | broker_produce_p999 | Raw data | 3 | > | 50 | - | ms | Hourly | |
| Creation Success Rate | broker_produce_success_rate | Raw data | 1 | < | - | 90 | % | Hourly | |
| Messages in the Dead Letter Queue | dlq_accumulation | Raw data | 3 | > | 0 | - | Count | Hourly | |
| Dead Letter Message Increase | dlq_increase | Raw data | 3 | > | 0 | - | Count | Hourly | |
| Topic Available Messages | topic_messages_remained | Raw data | 1 | > | 10000 | - | Count | Hourly | |
| Consumer Available Messages | consumer_messages_remained | Raw data | 1 | > | 10000 | - | Count | Hourly | |
| Socket Connections | socket_used | Raw data | 3 | > | 2500 | - | Count | Hourly | |
| Node Alive | rabbitmq_alive | Raw data | 1 | = | - | 0 | N/A | Hourly | |
| Disk Capacity Usage | rabbitmq_disk_usage | Raw data | 3 | > | 80 | 85 | % | Hourly | |
| CPU Usage | rabbitmq_cpu_usage | Raw data | 3 | - | > 80 | > 90% Or Decreased by 100% compared with the last period | % | Hourly | |
| Memory Usage | rabbitmq_memory_usage | Raw data | 3 | > | - | 30 | % | Hourly | |
| Memory High Watermark | rabbitmq_memory_high_watermark | Raw data | 1 | > | - | 0 | N/A | Hourly | |
| Disk High Watermark | rabbitmq_disk_insufficient | Raw data | 1 | > | - | 0 | N/A | Hourly | |
| Connection Usage | connections_usage | Raw data | 1 | > | - | 80 | % | Hourly | |
| Accumulated Messages | instance_accumulation | Raw data | 1 | > | 10000 | Increased by 50% compared with the last period | Count | Hourly | |
| Production Rate Limits | instance_produce_ratelimit_times | Raw data | 1 | >= | - | 1 | Count | Hourly | |
| Group Available Messages | group_accumulation | Raw data | 1 | > | 10000 | - | Count | Hourly | |
| Task Status | task_status | Raw data | 1 | = | 0 | - | N/A | Hourly | |
| Message Delay | message_delay | Raw data | 3 | > | 1000 | - | ms | Hourly | |
| Partitions | current_partitions | Raw data | 3 | > | 90 | - | Count | Hourly | |
| Accumulated Messages | group_msgs | Raw data | 1 | > | 10000 | - | Count | Hourly | |
| Queue Available Messages | queue_messages_ready | Raw data | 1 | > | 10000 | - | Count | Hourly | |
| Average Message Creation Processing Duration | broker_produce_mean | Raw data | 3 | > | - | 50 | ms | Hourly | |
| JVM Heap Memory Usage | broker_heap_usage | Raw data | 3 | > | 80 | 90 | % | Hourly | |
| Connections | broker_connections | Raw data | 1 | > | - | 4000 | Count | Hourly | |
| CPU Usage | broker_cpu_usage | Raw data | 3 | > | 80 | 90 | % | Hourly | |
| Network Bandwidth Usage | network_bandwidth_usage | Raw data | 3 | > | 70 | 80 | % | Hourly |
Relational Database Service
| Namespace | Metric Name | Metric ID | Statistic | Consecutive Triggers | Operator | Major Alarm Threshold | Critical Alarm Threshold | Unit | Frequency |
|---|---|---|---|---|---|---|---|---|---|
| SYS.RDS | CPU Usage | rds001_cpu_util | Raw data | 3 | >= | 80 | 90 | % | Hourly |
| Memory Usage | rds002_mem_util | Raw data | 3 | >= | 90 | 95 | % | Hourly | |
| Disk Usage | rds039_disk_util | Raw data | 3 | >= | 80 | 95 | % | Hourly | |
| Connection Usage | rds072_conn_usage | Raw data | 3 | >= | 80 | 90 | % | Hourly | |
| Real-Time Replication Delay | rds073_replication_delay | Raw data | 3 | >= | 300 | 600 | s | Hourly | |
| Active Connection Usage | rds_conn_active_usage | Raw data | 3 | >= | 80 | 95 | % | Hourly | |
| Stream Replication Status of Standby Instance or Read Replica | slave_replication_status | Raw data | 3 | = | - | 0 | Count | Hourly | |
| Replication Lag | rds046_replication_lag | Raw data | 3 | >= | 300000 | 600000 | ms | Hourly | |
| Connection Usage | rds083_conn_usage | Raw data | 3 | >= | 80 | 90 | % | Hourly |
Relational Database Service Cluster
| Namespace | Metric Name | Metric ID | Statistic | Consecutive Triggers | Operator | Major Alarm Threshold | Critical Alarm Threshold | Unit | Frequency |
|---|---|---|---|---|---|---|---|---|---|
| SYS.RDS_MYSQL_CLUSTER | Active Connection Usage | rds_conn_active_usage | Raw data | 3 | >= | 80 | 95 | % | Hourly |
| CPU Usage | rds001_cpu_util | Raw data | 3 | >= | 80 | 90 | % | Hourly | |
| Memory Usage | rds002_mem_util | Raw data | 3 | >= | 90 | 95 | % | Hourly | |
| Storage Space Usage | rds039_disk_util | Raw data | 3 | >= | 80 | 95 | % | Hourly | |
| Connection Usage | rds072_conn_usage | Raw data | 3 | >= | 80 | 90 | % | Hourly | |
| Real-Time Replication Delay | rds073_replication_delay | Raw data | 3 | >= | 300 | 600 | s | Hourly |
Content Delivery Network
| Namespace | Metric Name | Metric ID | Statistic | Consecutive Triggers | Operator | Major Alarm Threshold | Critical Alarm Threshold | Unit | Frequency |
|---|---|---|---|---|---|---|---|---|---|
| SYS.CDN | Bandwidth | bw | Raw data | 3 | Increase or decrease compared with last period | 10 | 20 | % | Hourly |
| Retrieval Failure Rate | bs_fail_rate | Raw data | 3 | > | 3 | 10 | % | Hourly | |
| Status Codes 4xx | http_code_4xx | Raw data | 3 | Increase compared with last period | 60 | 80 | % | Hourly | |
| 4xx Status Code Ratio | http_code_4xx_rate | Raw data | 3 | >= | 10 | 30 | % | Hourly | |
| Status Codes 5xx | http_code_5xx | Raw data | 3 | Increase compared with last period | 60 | 80 | % | Hourly | |
| 5xx Status Code Ratio | http_code_5xx_rate | Raw data | 3 | > | 1 | 5 | % | Hourly | |
| Traffic Hit Ratio | hit_flux_rate | Raw data | 3 | < | 80 | 50 | % | Hourly | |
| 5xx Origin Status Code Ratio | bs_http_code_5xx_rate | Raw data | 3 | > | 1 | 5 | % | Hourly |
Live
| Namespace | Metric Name | Metric ID | Statistic | Consecutive Triggers | Operator | Major Alarm Threshold | Critical Alarm Threshold | Unit | Frequency |
|---|---|---|---|---|---|---|---|---|---|
| SYS.LIVE | 5xx Status Code Proportion | http_5xx_proportion | Raw data | 1 | > | 0 | 1 | % | Hourly |
Data Warehouse Service
| Namespace | Metric Name | Metric ID | Statistic | Consecutive Triggers | Operator | Major Alarm Threshold | Critical Alarm Threshold | Unit | Frequency |
|---|---|---|---|---|---|---|---|---|---|
| SYS.DWS | CPU Usage | dws010_cpu_usage | Raw data | 3 | > | 85 | 90 | % | Daily |
| Memory Usage | dws011_mem_usage | Raw data | 3 | > | 90 | 95 | % | Daily | |
| Disk Usage | dws015_disk_usage | Raw data | 3 | > | 80 | 90 | % | Daily | |
| Disk Read Throughput | dws018_disk_read_throughput | Raw data | 5 | > | - | 300000000 | Byte/s | 6 hours | |
| Disk Write Throughput | dws019_disk_write_throughput | Raw data | 5 | > | - | 300000000 | Byte/s | 6 hours |
Data Replication Service
| Namespace | Metric Name | Metric ID | Statistic | Consecutive Triggers | Operator | Major Alarm Threshold | Critical Alarm Threshold | Unit | Frequency |
|---|---|---|---|---|---|---|---|---|---|
| SYS.DRS | CPU Usage | cpu_util | Raw data | 3 | > | 80 | 90 | % | Hourly |
| Memory Usage | mem_util | Raw data | 3 | > | 85 | 90 | % | Hourly | |
| Storage Space Usage | disk_util | Raw data | 3 | > | 80 | 90 | % | Hourly | |
| Source Database WAL Extract Lag | extract_latency | Raw data | 3 | > | 300000 | 600000 | ms | Hourly | |
| Data Synchronization Latency | apply_latency | Raw data | 3 | > | 300000 | 600000 | ms | Hourly | |
| Synchronization Status | apply_current_state | Raw data | 3 | = | - | 10 | N/A | Hourly | |
| Task Status | apply_job_status | Raw data | 3 | = | - | 1 | N/A | Hourly |
Database Security Service
| Namespace | Metric Name | Metric ID | Statistic | Consecutive Triggers | Operator | Major Alarm Threshold | Critical Alarm Threshold | Unit | Frequency |
|---|---|---|---|---|---|---|---|---|---|
| SYS.DBSS | CPU Usage | cpu_util | Raw data | 3 | > | 80 | 85 | % | Hourly |
| Memory Usage | mem_util | Raw data | 3 | > | 80 | 85 | % | Hourly | |
| Disk Usage | disk_util | Raw data | 3 | > | 80 | 85 | % | Hourly |
Database Proxy
| Namespace | Metric Name | Metric ID | Statistic | Consecutive Triggers | Operator | Major Alarm Threshold | Critical Alarm Threshold | Unit | Frequency |
|---|---|---|---|---|---|---|---|---|---|
| SYS.DBPROXY | CPU Usage | rds001_cpu_util | Raw data | 3 | > | 80 | 90 | % | Hourly |
| Memory Usage | rds002_mem_util | Raw data | 3 | > | 90 | 95 | % | Hourly | |
| Intranet Outbound Bandwidth Usage (%) | l4_out_bps_usage | Raw data | 2 | > | 90 | 95 | % | Hourly | |
| Intranet Inbound Bandwidth Usage (%) | l4_in_bps_usage | Raw data | 2 | > | 90 | 95 | % | Hourly | |
| Abnormal ELB Backend Proxy Nodes | m9_abnormal_servers | Raw data | 1 | > | - | 0 | Count | Hourly |
Document Database Service
| Namespace | Metric Name | Metric ID | Statistic | Consecutive Triggers | Operator | Major Alarm Threshold | Critical Alarm Threshold | Unit | Frequency |
|---|---|---|---|---|---|---|---|---|---|
| SYS.DDS | Replication Lag | mongo026_repl_lag | Raw data | 3 | >= | 300 | 600 | s | Hourly |
| CPU Usage | mongo031_cpu_usage | Raw data | 3 | >= | 80 | 98 | % | Hourly | |
| Memory Usage | mongo032_mem_usage | Raw data | 3 | >= | 90 | 98 | % | Hourly | |
| Storage Space Usage | mongo035_disk_usage | Raw data | 3 | >= | 80 | 95 | % | Hourly | |
| Disk Read Time | mongo039_avg_disk_sec_per_read | Raw data | 3 | >= | 0.05 | 0.1 | s | Hourly | |
| Disk Write Time | mongo040_avg_disk_sec_per_write | Raw data | 3 | >= | 0.05 | 0.1 | s | Hourly | |
| Percentage of Active Node Connections | mongo007_connections_usage | Raw data | 3 | >= | 80 | 95 | % | Hourly | |
| WiredTiger Cache Percentage | mongo054_wt_cache_used_percent | Raw data | 3 | >= | 85 | 95 | % | Hourly | |
| WiredTiger Dirty Data Cache Percentage | mongo055_wt_cache_dirty_percent | Raw data | 3 | >= | 20 | 25 | % | Hourly |
Virtual Private Cloud
| Namespace | Metric Name | Metric ID | Statistic | Consecutive Triggers | Operator | Major Alarm Threshold | Critical Alarm Threshold | Unit | Frequency |
|---|---|---|---|---|---|---|---|---|---|
| SYS.VPC | Outbound Bandwidth Usage | upstream_bandwidth_usage | Raw data | 3 | > | - | 80 | % | Hourly |
| Inbound Bandwidth Usage | downstream_bandwidth_usage | Raw data | 3 | > | - | 80 | % | Hourly | |
| Outbound Bandwidth Usage | upstream_bandwidth_usage | Raw data | 3 | Increase or decrease compared with last period | 20 | - | % | Hourly | |
| Inbound Bandwidth Usage | downstream_bandwidth_usage | Raw data | 3 | Increase or decrease compared with last period | 20 | - | % | Hourly |
Cloud Firewall
| Namespace | Metric Name | Metric ID | Statistic | Consecutive Triggers | Operator | Major Alarm Threshold | Critical Alarm Threshold | Unit | Frequency |
|---|---|---|---|---|---|---|---|---|---|
| SYS.CFW | Protection Bandwidth Usage Rate | protection_bandwidth_usage | Raw data | 3 | > | 85 | 95 | % | Hourly |
| Internet Boundary Protection Bandwidth Usage (%) | internet_protection_bandwidth_usage_rate | Raw data | 3 | > | 85 | 95 | % | Hourly | |
| Inter-VPC Protection Bandwidth Usage (%) | vpc_protection_bandwidth_usage_rate | Raw data | 3 | > | 85 | 95 | % | Hourly |
Cloud Connect
| Namespace | Metric Name | Metric ID | Statistic | Consecutive Triggers | Operator | Major Alarm Threshold | Critical Alarm Threshold | Unit | Frequency |
|---|---|---|---|---|---|---|---|---|---|
| SYS.CC | Network Bandwidth Usage | network_bandwidth_usage | Raw data | 3 | > | - | 80 | % | Hourly |
TaurusDB
| Namespace | Metric Name | Metric ID | Statistic | Consecutive Triggers | Operator | Major Alarm Threshold | Critical Alarm Threshold | Unit | Frequency |
|---|---|---|---|---|---|---|---|---|---|
| SYS.GAUSSDB | CPU Usage | gaussdb_mysql001_cpu_util | Raw data | 3 | > | 80 | 90 | % | Hourly |
| Memory Usage | gaussdb_mysql002_mem_util | Raw data | 3 | > | - | 90 | % | Hourly | |
| Connection Usage | gaussdb_mysql072_conn_usage | Raw data | 3 | > | 80 | 90 | % | Hourly | |
| Data Disk Usage | gaussdb_mysql113_data_disk_used_ratio | Raw data | 3 | > | 80 | 90 | % | Hourly |
Cloud Search Service
| Namespace | Metric Name | Metric ID | Statistic | Consecutive Triggers | Operator | Major Alarm Threshold | Critical Alarm Threshold | Unit | Frequency |
|---|---|---|---|---|---|---|---|---|---|
| SYS.ES | Max. Disk Usage | disk_util | Raw data | 5 | >= | 85 | 90 | % | Hourly |
| Cluster Health Status | status | Raw data | 5 | >= | 1 | 2 | N/A | Hourly | |
| Max. JVM Heap Usage | max_jvm_heap_usage | Raw data | 1 | > | 80 | 85 | % | Hourly | |
| Max. CPU Usage | max_cpu_usage | Raw data | 2 | > | 80 | 85 | % | Hourly | |
| Nodes | nodes_count | Raw data | 3 | Decrease compared with previous period | - | 10 | % | Hourly | |
| Tasks in Write Queue | sum_thread_pool_write_queue | Raw data | 5 | >= | 500 | 1000 | N/A | Hourly | |
| Tasks in Search Queue | sum_thread_pool_search_queue | Raw data | 5 | >= | 500 | 800 | N/A | Hourly | |
| Rejected Tasks in Write Queue | sum_thread_pool_write_rejected | Raw data | 5 | >= | 10 | 20 | N/A | Hourly | |
| Rejected Tasks in Search Queue | sum_thread_pool_search_rejected | Raw data | 5 | >= | 10 | 20 | N/A | Hourly | |
| Max. Task Runtime | task_max_running_time | Raw data | 1 | >= | - | 60000 | ms | Hourly |
Direct Connect
| Namespace | Metric Name | Metric ID | Statistic | Consecutive Triggers | Operator | Major Alarm Threshold | Critical Alarm Threshold | Unit | Frequency |
|---|---|---|---|---|---|---|---|---|---|
| SYS.DCAAS | Port Status | network_status | Raw data | 1 | != | - | 1 | N/A | Every 5 minutes |
| Inbound Error Packets | in_errors | Raw data | 1 | > | - | 0 | Packet | Every 5 minutes | |
| Latency | latency | Raw data | 3 | Increase compared with last period | - | 20 | % | Hourly | |
| Packet Loss Rate | packet_loss_rate | Raw data | 3 | > | 5 | 10 | % | Hourly | |
| IPv4 Peer Status | bgp_peer_status_v4 | Raw data | 1 | != | - | 1 | N/A | Hourly | |
| IPv6 Peer Status | bgp_peer_status_v6 | Raw data | 1 | != | - | 1 | N/A | Hourly |
Feedback
Was this page helpful?
Provide feedbackThank you very much for your feedback. We will continue working to improve the documentation.See the reply and handling status in My Cloud VOC.
For any further questions, feel free to contact us through the chatbot.
Chatbot