Checking Event Inspection Data on AOM
AOM regularly inspects application services for which the intelligent patrol function has been enabled, monitors service quality based on key metrics (such as the average RT and error rate) of historical data, and enables global analysis of problems.
Description
AOM dynamically determines upper limits based on the historical data of applications and checks whether the recent data is abnormal.
- Dynamically determines upper limits based on the historical 3-hour data of applications and checks whether the data in the last 10 minutes is abnormal. The following event types are supported:
- Service Avg. RT Sharply Increases
Based on the historical 3-hour data of applications, AOM determines whether the average RT of application services increases sharply in the last 10 minutes.
- Top N API Avg. RT Sharply Increases
By default, the five APIs with the most traffic are checked. Based on the historical 3-hour data of APIs, AOM determines whether the average RT of the top 5 APIs sharply increases in the last 10 minutes.
- Service Error Rate Sharply Increases
Based on the historical 3-hour data of applications, AOM determines whether the error rate of application services increases sharply in the last 10 minutes.
- Top N API Error Rate Sharply Increases
By default, the five APIs with the most traffic are checked. Based on the historical 3-hour data of APIs, AOM determines whether the error rate of the top 5 APIs sharply increases in the last 10 minutes.
- Service Avg. RT Sharply Increases
- Dynamically determines upper limits based on the historical 1-hour data of applications and checks whether the data in the last 15 minutes is abnormal.
- Top N API Traffic Uneven: By default, the five APIs with the most traffic (number of calls) are checked. Based on the historical 1-hour data of APIs, AOM determines whether the traffic of the top 5 APIs exceeds or falls below the preset upper or lower limit in the last 15 minutes.
Procedure
- Log in to the AOM 2.0 console.
- In the navigation pane, choose . The Intelligent Patrol page is displayed.
- Events can be displayed in List or Card mode. You can view events by application, event type, event description, and event status. The event statistics filtered based on criteria are displayed in the event overview area. The following shows an example.
- Event overview
In the graph area, perform the following operations if needed:
- In the upper left corner of the graph, view the total number of abnormal events detected during inspection in the specified period.
- Move the pointer to the bar graph to view the number of events of each type at a specific time point.
- Click a legend above the bar graph to hide or display a certain type of events.
- In the search box, enter a keyword to filter events.
- List mode
The abnormal events detected by intelligent patrol in the specified period are displayed in a list. Each event contains the following information:
- Event Type: type of an event.
- Application: application to which the event belongs.
- Root Cause Component: component that causes the event.
- Description: describes the component and interface where the event occurs.
- Triggered: time when an exception first occurs.
- Duration: the period for which the exception lasts.
- Status: status of the event during intelligent patrol.
- Card mode
The abnormal events detected by intelligent patrol in the specified period are displayed on cards. Each card contains the following information:
- Event Type: type of an event.
- Application: application to which the event belongs.
- Root Cause Component: component that causes the event.
- Description: describes the component and interface where the event occurs.
- Triggered: time when an exception first occurs.
- Duration: the period for which the exception lasts.
- Status: status of the event during intelligent patrol.
- Event overview
- View the details of an event. Click an event type name in the list or on the card to go to the event details page.
On the event details page, graphs about key metrics such as RT and error rate are displayed, showing the duration for which an exception lasts, time when the exception first occurs, and upper limit. (You can click the component, environment, or API name under Problem Description on the event details page to go to the corresponding details page. Only the AP-Singapore region supports the redirection to the component details page.)
Feedback
Was this page helpful?
Provide feedbackThank you very much for your feedback. We will continue working to improve the documentation.See the reply and handling status in My Cloud VOC.
For any further questions, feel free to contact us through the chatbot.
Chatbot