Overview
COC fault management provides you with the capabilities of quick fault demarcation, locating, and rectification. It supports ingestion of alarms from multiple data sources. COC aggregates raw alarms for alarm noise reduction and then convert corresponding alarms into incidents and leave others as aggregated alarms. Faults reported by the alarms or incidents will be quickly demarcated through the application topology diagnosis tool or war rooms, and then be swiftly rectified based on online response plans with the MTTR shortened. All faults and their handling processes will be reviewed for service improvement. In addition, with this feature, COC continuously accumulates the fault management O&M knowledge base and improves the risk resistance capability.
Fault Management Introduction
Core Features
- Data Sources: Raw alarms from multiple monitoring platforms are ingested into COC for central management. Huawei Cloud Eye, AOM, APM, LTS, Alibaba CloudMonitor, Alibaba Simple Log Service (SLS), Prometheus, Grafana, Zabbix, and user-defined service monitoring systems (connected through open APIs) can be interconnected.
- Incident Forwarding Rules: This module is used after certain alarm data sources are interconnected. You can set trigger conditions and rules to convert raw alarms into aggregated alarms or incident tickets on COC. You can also assign owners and preset response plans for aggregated alarms or incident tickets.
- Alarms: This module displays raw alarms and aggregated alarms. You can perform operations on aggregated alarms, including alarm clearance, alarm-to-incident operations, and response plan execution.
- Incidents: This module manages incident tickets throughout the lifecycle, including incident ticket creation, acceptance, rejection, forwarding, processing, escalation, de-escalation, and war room initiation.
- War Rooms: This module applies when major faults occur and different roles need to be quickly gathered to locate and rectify the faults. This module displays affected applications, related alarms, incidents, change information, and rectification progress notifications. You can execute response plans, diagnose applications, and start groups of third-party OA software.
- Issues: This module displays issues detected, recorded, and resolved during the use of software products, such as functional defects and poor performance.
- Improvement Tickets: This module displays product, O&M, or management improvement items identified during troubleshooting. You can trace and close them through online improvement tickets.
Feedback
Was this page helpful?
Provide feedbackThank you very much for your feedback. We will continue working to improve the documentation.See the reply and handling status in My Cloud VOC.
For any further questions, feel free to contact us through the chatbot.
Chatbot