Managing Failure Modes
A failure mode refers to a specific type of problem or failure status that may occur during application running. Build a rich failure mode library and formulate corresponding prevention and recovery measures to help design a more highly available application system. By identifying potential faults, you can perform routine drills to verify whether the fault recovery measures and fault impacts meet the expectations and prepare for better response to various challenges. You can analyze the possible fault points of an application, create a failure mode by describing the fault occurrence conditions, fault symptoms, and customer impacts, and apply the failure mode to routine chaos drills.
Relationship Between Failure Modes and Drill Tasks
A failure mode and a drill task are closely linked and building upon each other in the chaos drill system. They form a closed loop for risk identification through fault injection and capability verification through task execution.
Failure modes focus on risk assessment of cloud applications. By systematically evaluating the application architecture, dependency relationship, and potential weak points, the system can precisely identify risk scenarios that could lead to service anomalies (such as node failures, network latency, resource exhaustion). Failure modes are the core premise and basis for conducting chaos drills.
Drill tasks are the carriers for failure modes to be implemented. Based on these failure modes, individual or interconnected fault scenarios are designed and simulated. Fault injection tools (such as server outage simulation and traffic congestion injection) are then employed to replicate the corresponding risks in a controlled environment. Finally, the application's fault tolerance, auto-recovery efficiency, and the effectiveness of contingency plans are validated. This process enables the transformation of risk identification into tangible capability verification.
Precautions
Verify that the enterprise project, application, event level, and scenario category of the failure mode are correct.
Creating a Failure Mode
- Log in to COC.
- In the navigation pane, choose Resilience Center > Chaos Drills.
- Click Failure Modes. On the displayed page, click Create Failure Mode.
- On the displayed page, set parameters for creating a failure mode.
Table 1 Parameters for creating a failure mode Parameter
Description
Failure Mode
Customize a failure mode name.
The name can contain a maximum of 64 characters, including letters, digits, hyphens (-), underscores (_), and spaces. It cannot start or end with a space.
Scenario Category
The options are Node, Cluster, Network, DR, Container, and Services and Data.
- Node: The CPU or memory of the host is overloaded, or the process is faulty. As a result, services are abnormal, for example, the CPU or memory is overloaded, or the process status is abnormal.
- Cluster: Simulate abnormal scenarios by increasing the pressure or performing an active/standby cluster switchover, for example, increasing the pressure of the container cluster and performing an active/standby switchover in the database cluster.
- Network: Inject network faults to hosts or clusters to verify the DR capability of your service. Such faults include packet loss, network latency, and intermittent disconnection at the link layer.
- DR: Simulate inter-region network exceptions or service unavailability in a single region to verify the self-recovery capabilities of services.
- Container: Simulate process and resource faults and network attacks on container instances, such as CPU and memory pressure increase, network attacks, system OOM, or process killing.
- Services and Data: Simulate service exceptions caused by database or file exceptions. Such faults include packet loss, network latency, and intermittent disconnection at the link layer.
Incident Level
The options are P1, P2, P3, P4, and P5.
P1 incidents are the most critical, while P5 incidents are the least severe.
Source
The options are Failure modes detected proactively and Existing failure modes.
- Failure modes detected proactively: Risks are analyzed proactively in the application architecture and running environment to form a failure mode.
- Existing failure modes: A failure mode is formed based on the analysis of existing faults and incidents.
Alarm ID
(Optional) ID of the alarm that is triggered when a fault occurs.
Attack Scenario
(Optional) Select an attack scenario.
A maximum of 10 attack scenarios can be selected.
Enterprise Project
Select the enterprise project to which the failure mode resource belongs.
- Select an existing enterprise project.
- Click Create to create an enterprise project and select it. For details, see Creating an Enterprise Project.
After a failure mode is created, the enterprise project cannot be changed.
Application
Select the application to which the drill target belongs.
- Select an existing application.
- Click Create to create an application and select it. For details, see Creating an Application.
Contingency Plan Available
The options are Yes and No.
Contingency Plan
This parameter is mandatory when Contingency Plan Available is set to Yes.
- Select an existing contingency plan.
- Click Create Contingency Plan to create a contingency plan and select it. For details, see Creating a Contingency Plan.
Occurrence Conditions
Enter the conditions under which the fault may occur.
The value can contain a maximum of 1,024 characters.
Fault Symptom
Enter the possible service symptom when the fault occurs.
The value can contain a maximum of 1,024 characters.
Impact on Customer
(Optional) Enter the impact of the fault on customers.
The value can contain a maximum of 1,024 characters.
- Click OK.
After the failure mode is created, you can view it in the custom failure mode list and plan the drill.
Cloning a Failure Mode
You can clone a failure mode to quickly generate a custom failure mode, which greatly reduces the cost of creating a new one.
Only the failure mode in the Preset Failure Mode Cases page can be cloned.
- Log in to COC.
- In the navigation pane, choose Resilience Center > Chaos Drills.
- Choose Failure Mode > Preset Failure Mode Cases. On the displayed page, locate the target failure mode and click Clone in the Operation column.
- Adjust the failure mode information based on the service scenario.
Table 2 Parameters for cloning a failure mode Parameter
Description
Failure Mode
Customize a failure mode name.
The name can contain a maximum of 64 characters, including letters, digits, hyphens (-), underscores (_), and spaces. It cannot start or end with a space.
Scenario Category
The options are Node, Cluster, Network, DR, Container, and Services and Data.
- Node: The CPU or memory of the host is overloaded, or the process is faulty. As a result, services are abnormal, for example, the CPU or memory is overloaded, or the process status is abnormal.
- Cluster: Simulate abnormal scenarios by increasing the pressure or performing an active/standby cluster switchover, for example, increasing the pressure of the container cluster and performing an active/standby switchover in the database cluster.
- Network: Inject network faults to hosts or clusters to verify the DR capability of your service. Such faults include packet loss, network latency, and intermittent disconnection at the link layer.
- DR: Simulate inter-region network exceptions or service unavailability in a single region to verify the self-recovery capabilities of services.
- Container: Simulate process and resource faults and network attacks on container instances, such as CPU and memory pressure increase, network attacks, system OOM, or process killing.
- Services and Data: Simulate service exceptions caused by database or file exceptions. Such faults include packet loss, network latency, and intermittent disconnection at the link layer.
Incident Level
The options are P1, P2, P3, P4, and P5.
P1 incidents are the most critical, while P5 incidents are the least severe.
Source
The options are Failure modes detected proactively and Existing failure modes.
- Failure modes detected proactively: Risks are analyzed proactively in the application architecture and running environment to form a failure mode.
- Existing failure modes: A failure mode is formed based on the analysis of existing faults and incidents.
Alarm ID
(Optional) ID of the alarm that is triggered when a fault occurs.
Attack Scenario
(Optional) Select an attack scenario.
A maximum of 10 attack scenarios can be selected.
Enterprise Project
Select the enterprise project to which the failure mode resource belongs.
After a failure mode is cloned, the enterprise project cannot be changed.
Application
Select the application to which the drill target belongs.
Contingency Plan Available
The options are Yes and No.
Contingency Plan
This parameter is mandatory when Contingency Plan Available is set to Yes.
Select a contingency plan. If no best-fit contingency plans are available, create one. For details, see Creating a Contingency Plan.
Occurrence Conditions
Enter the conditions under which the fault may occur.
The value can contain 1 to 1,024 characters.
Fault Symptom
Enter the possible service symptom when the fault occurs.
The value can contain 1 to 1,024 characters.
Impact on Customer
Describe the impact of the fault on customers.
The value can contain 0 to 1,024 characters.
- Click OK.
After the failure mode is cloned, you can view the cloned failure mode on the Custom Failure Modes page.
Modifying a Custom Failure Mode
Only custom fault modes can be modified.
- Log in to COC.
- In the navigation pane, choose Resilience Center > Chaos Drills.
- Choose Failure mode > Custom Failure Modes. On the displayed page, locate the target failure mode and click Modify in the Operation column.
- Modify the failure mode information as required.
Table 3 Failure mode parameters Parameter
Description
Failure Mode
Customize a failure mode name.
The name can contain a maximum of 64 characters, including letters, digits, hyphens (-), underscores (_), and spaces. It cannot start or end with a space.
Scenario Category
The options are Node, Cluster, Network, DR, Container, and Services and Data.
- Node: The CPU or memory of the host is overloaded, or the process is faulty. As a result, services are abnormal, for example, the CPU or memory is overloaded, or the process status is abnormal.
- Cluster: Simulate abnormal scenarios by increasing the pressure or performing an active/standby cluster switchover, for example, increasing the pressure of the container cluster and performing an active/standby switchover in the database cluster.
- Network: Inject network faults to hosts or clusters to verify the DR capability of your service. Such faults include packet loss, network latency, and intermittent disconnection at the link layer.
- DR: Simulate inter-region network exceptions or service unavailability in a single region to verify the self-recovery capabilities of services.
- Container: Simulate process and resource faults and network attacks on container instances, such as CPU and memory pressure increase, network attacks, system OOM, or process killing.
- Services and Data: Simulate service exceptions caused by database or file exceptions. Such faults include packet loss, network latency, and intermittent disconnection at the link layer.
Incident Level
The options are P1, P2, P3, P4, and P5.
P1 incidents are the most critical, while P5 incidents are the least severe.
Source
The options are Failure modes detected proactively and Existing failure modes.
- Failure modes detected proactively: Risks are analyzed proactively in the application architecture and running environment to form a failure mode.
- Existing failure modes: A failure mode is formed based on the analysis of existing faults and incidents.
Alarm ID
(Optional) ID of the alarm that is triggered when a fault occurs.
Attack Scenario
(Optional) Select an attack scenario.
A maximum of 10 attack scenarios can be selected.
Enterprise Project
After a failure mode is created, the enterprise project cannot be changed.
Application
Select the application to which the drill target belongs.
Contingency Plan Available
The options are Yes and No.
Contingency Plan
This parameter is mandatory when Contingency Plan Available is set to Yes.
Select a contingency plan. If no best-fit contingency plans are available, create one. For details, see Creating a Contingency Plan.
Occurrence Conditions
Enter the conditions under which the fault may occur.
The value can contain 1 to 1,024 characters.
Fault Symptom
Enter the possible service symptom when the fault occurs.
The value can contain 1 to 1,024 characters.
Impact on Customer
Describe the impact of the fault on customers.
The value can contain 0 to 1,024 characters.
- Click OK.
Deleting a Custom Failure Mode
Only custom fault modes can be deleted. The fault mode has been associated with a drill task or drill plan and cannot be deleted.
- Log in to COC.
- In the navigation pane, choose Resilience Center > Chaos Drills.
- Choose Failure mode > Custom Failure Modes. On the displayed page, locate the target failure mode and click Delete in the Operation column.
- In the displayed dialog box, click OK.
After the failure mode is deleted, it will no longer be displayed in the custom failure mode list.
Helpful Links
Feedback
Was this page helpful?
Provide feedbackThank you very much for your feedback. We will continue working to improve the documentation.See the reply and handling status in My Cloud VOC.
For any further questions, feel free to contact us through the chatbot.
Chatbot