Updated on 2026-06-23 GMT+08:00

Creating a Drill Task

You can simulate software or hardware faults to test the system's fault rectification capability by using a drill task. You can manage chaos drill tasks, view drill records, and create drill tasks. You can set basic information, add attack task groups, and select attack tasks and attack scenarios for a drill task. In addition, a drill task involves monitoring task configuration and post-drill review and improvement. This ensures that an excellent optimization policy can be applied when the system is under various pressures.

Operation

Procedure Description

Figure 1 Drill process
  1. Create a drill task: To evaluate the system's recovery capability under simulated hardware or software failures, configure relevant fault scenarios, target applications, and monitoring rules based on business requirements—whether ad hoc, custom-tailored, or periodic and standardized—to establish the foundational structure for a chaos drill.
  2. Start a drill task: Once created, the chaos drill task is formally executed. The system simulates the predefined faults and progresses through the drill workflow, while supporting flexible runtime controls such as termination, retry, or skip in response to real-time conditions (such as anomalies or timeouts).
  3. Create a drill report: After the drill task completes, a comprehensive report is produced, summarizing the entire drill process, outcome data (including recovery duration), and identified improvement items. This report is exportable and traceable, providing actionable insights for system optimization and refinement of emergency response strategies.
  4. View drill history: You can retrieve key details of completed chaos drills, including baseline configuration, execution status, fault injection specifics, and monitoring metrics, enabling full visibility into drill implementation and the system's actual performance during failure scenarios.

Automatic Task Termination Mechanism

  • Automatic termination upon timeout: If a drill task fails and you do not manually close the task within 48 hours, the system automatically terminates it.
  • Automatic termination upon exceptions: During the drill execution, if a pod exception (for example, the pod has been deleted) is detected or a resource O&M ticket is manually closed, the system automatically terminates the current task immediately.

Creating a Drill Task

You can create a drill task either directly or using a template. Based on your service scenario, drill complexity, and standardization needs, you can flexibly choose the most suitable approach to ensure efficient implementation of drill tasks and enhance both the standardization and execution efficiency of emergency drills.

Create a drill task directly: This approach is suitable for ad hoc drill requirements, customized scenarios, or first-time drills where no existing template is available—for example, targeted emergency drills in response to an unexpected fault or temporary service process validation exercises.

Create a drill task using a template: This approach is suitable for periodic, repeatable drills, standardized process exercises, or scenarios requiring consistent practices across multiple teams—for example, monthly DR switchover drills, quarterly security incident response drills, and routine cross-departmental process drills.

  1. Log in to COC.
  2. In the navigation pane, choose Resilience Center > Chaos Drills.
  3. Click Drill Tasks.
  4. On the displayed page, click Create Task.

    You can also use the drill plan ticket accepting function to access the page for creating a drill task. For details, see Managing Drill Plans.

  5. Set the basic information.

    Table 1 Basic information parameters

    Parameter

    Description

    Example Value

    Drill Task

    Name of the drill task. Set it according to the naming rules.

    The name can contain a maximum of 64 characters, including letters, digits, hyphens (-), underscores (_), and spaces. It cannot start or end with a space.

    test-drill

    Expected Recovery Duration (Minutes)

    Expected time from the fault occurrence to the fault rectification, in minutes.

    Expected time for the application to recover to the normal state after fault injection or when the contingency plan is executed. This time does not affect the drill task.

    3

  6. Click Create Attack Task.

    By default, there is one attack task group. You can click Create Task Group to add a task group. Then, click Create Attack Task to add another attack task.
    • Tasks in different task groups are executed in serial mode, and tasks in the same task group are executed in parallel mode.
    • Currently, multiple fault injection operations on the same resource in a task group are not supported.
    • To add an existing task, click Select from Existing, select the existing task, and click OK.
    • To add a new attack task, perform the following steps.
      1. Set the attack target by referring to Table 2.
        Table 2 Parameters for adding an attack task

        Parameter

        Description

        Example Value

        Cloud Service Vendor

        Select a cloud service vendor type.

        You can choose Huawei Cloud, On-premises IDCs, or Alibaba Cloud.

        Huawei Cloud

        Source of Attack Target

        Select the source of the target instance.

        Elastic Cloud Server (ECS)

        Attack Task

        Customize the name of the attack task based on the naming rule.

        The preset name is the attack target type plus the creation timestamp. For example, if the attack target source is ECS and the creation time is 14:36:08 on December 18, 2025, the name is automatically generated as ECS_20251218143608.

        test-attacktask

        Attack Target

        Select the target instance.

        You can filter attack targets by application.

        You can select attack targets by selecting instances, pods, or a specified number of targets if CCE instances are used.

        -

      1. Click Next and select an attack scenario.
        For details about attack scenarios, see Appendix: Attack Scenarios.
        Table 3 Parameters for selecting an attack scenario

        Parameter

        Description

        Example Value

        Attack Type

        Select an attack type based on the attack scenario.

        Host resources

        Attack Scenario

        Select an attack scenario.

        CPU usage increase

        Attack Parameters

        Configure attack parameters based on attack scenarios.

        • CPU usage (%): 80
        • Fault duration (s): 60
      2. Click Next.
      3. (Optional) Set Configure Monitoring Tasks.
        Table 4 Parameters for configuring a monitoring task

        Parameter

        Description

        Steady-State Metrics

        Select the target resource, performance metric, lower limit, and upper limit from the drop-down lists one by one.

        If a service can perform well and stably when a performance monitoring metric is set to a certain value range, this metric is called a steady-state metric. If this metric value is not within that value range before the drill execution, the drill will be canceled.

        Metric

        Select the target resource, monitoring metric, lower limit, and upper limit from the drop-down lists one by one.

        These service metrics monitor the corresponding service data during fault drills. If the value of such a metric is within the allowed value range, the service is normal. Otherwise, you can determine whether to stop a drill.

        Automatic Rollback

        Select whether to enable automatic rollback.

        Fault injection is automatically rolled back and restored to the status before fault injection. Automatic rollback cannot be configured for some disruptors in fault drills that do not support fault termination.

        If the value of a steady-state metric is not within the stable value range during a drill, the corresponding fault injection automatically stops after automatic rollback is enabled.

        For details about metrics supported by each instance, see Cloud Product Metrics.

      4. Click Finish.

  7. Click OK. The drill task is created and the task status is to be drilled.

    If you click Save as Draft, the task status is draft, which does not allow you to start the drill task.

  1. Log in to COC.
  2. In the navigation pane, choose Resilience Center > Drill Templates.
  3. Create a drill task in either of the following ways:

  4. Set the basic information.

    Table 5 Basic information parameters

    Parameter

    Description

    Example Value

    Drill Task

    Name of the drill task. Set it according to the naming rules.

    test-drill

    Expected Recovery Duration (Minutes)

    Expected time from the fault occurrence to the fault rectification, in minutes.

    Expected time for the application to recover to the normal state after fault injection or when the contingency plan is executed. This time does not affect the drill task.

    3

  5. In the task group that contains scenario and parameters, locate the scenario and add an attack target for the task.

    1. Click Select under Attack Target.
    2. Cloud Service Provider and Source of Attack Target are selected based on the preset value of the scenario.
    3. In the Attack Target table, the instances that do not support the current disruptor are dimmed. After you select an existing task, change the cloud server provider, or change the source of the attack target, the preset disruptor information will be cleared. For details, see Creating a Drill Task Directly.
    4. After you select an attack target and click Next, the corresponding disruptor is selected based on the target scenario. The preset value of the attack parameter is the data in the template.
    5. (Optional) Set Configure Monitoring Tasks.
      Table 6 Parameters for configuring a monitoring task

      Parameter

      Description

      Steady-State Metrics

      Select the target resource, performance metric, lower limit, and upper limit from the drop-down lists one by one.

      If a service can perform well and stably when a performance monitoring metric is set to a certain value range, this metric is called a steady-state metric. If this metric value is not within that value range before the drill execution, the drill will be canceled. If the value of a steady-state metric is not within the stable value range during a drill, the corresponding fault injection automatically stops after automatic rollback is enabled.

      Metric

      Select the target resource, monitoring metric, lower limit, and upper limit from the drop-down lists one by one.

      These service metrics monitor the corresponding service data during fault drills. If the value of such a metric is within the allowed value range, the service is normal. Otherwise, you can determine whether to stop a drill.

      Automatic Rollback

      Select whether to enable automatic rollback.

      Fault injection is automatically rolled back and restored to the status before fault injection. Automatic rollback cannot be configured for some disruptors in fault drills that do not support fault termination.

    6. Click Finish.

  6. (Optional) Click Create Attack Task.

    Attack tasks are preset in the template. You can add an attack task as required. Click Create Task Group. Then, click Create Attack Task to add another attack task.
    • Tasks in different task groups are executed in serial mode, and tasks in the same task group are executed in parallel mode.
    • Currently, multiple fault injection operations on the same resource in a task group are not supported.
    • To add an existing task, click Select from Existing, select the existing task, and click OK.
    • To add a new attack task, perform the following steps.
      1. Set the attack target.
        Table 7 Parameters for adding an attack task

        Parameter

        Description

        Example Value

        Cloud Service Vendor

        Select a cloud service vendor type.

        You can choose Huawei Cloud, On-premises IDCs, or Alibaba Cloud.

        Huawei Cloud

        Source of Attack Target

        Select the source of the target instance.

        Elastic Cloud Server (ECS)

        Attack Task

        Customize the name of the attack task based on the naming rule.

        The preset name is the attack target type plus the creation timestamp. For example, if the attack target source is ECS and the creation time is 14:36:08 on December 18, 2025, the name is automatically generated as ECS_20251218143608.

        test-attacktask

        Attack Target

        Select the target instance.

        You can filter attack targets by application.

        You can select attack targets by selecting instances, pods, or a specified number of targets if CCE instances are used.

        -

      2. Click Next.
      3. Set parameters for selecting an attack scenario.
        For details, see Appendix: Attack Scenarios.
        Table 8 Parameters for selecting an attack scenario

        Parameter

        Description

        Example Value

        Attack Type

        Select an attack type based on the attack scenario.

        Host resources

        Attack Scenario

        Customize the name of the attack task based on the naming rule.

        CPU usage increase

        Attack Parameters

        Configure attack parameters based on attack scenarios.

        • CPU usage (%): 80
        • Fault duration (s): 60
      4. Click Next.
      5. (Optional) Set Configure Monitoring Tasks.
        Table 9 Parameters for configuring a monitoring task

        Parameter

        Description

        Steady-State Metrics

        Select the target resource, performance metric, lower limit, and upper limit from the drop-down lists one by one.

        If a service can perform well and stably when a performance monitoring metric is set to a certain value range, this metric is called a steady-state metric. If this metric value is not within that value range before the drill execution, the drill will be canceled. If the value of a steady-state metric is not within the stable value range during a drill, the corresponding fault injection automatically stops after automatic rollback is enabled.

        Metric

        Select the target resource, monitoring metric, lower limit, and upper limit from the drop-down lists one by one.

        These service metrics monitor the corresponding service data during fault drills. If the value of such a metric is within the allowed value range, the service is normal. Otherwise, you can determine whether to stop a drill.

        Automatic Rollback

        Select whether to enable automatic rollback.

        Fault injection is automatically rolled back and restored to the status before fault injection. Automatic rollback cannot be configured for some disruptors in fault drills that do not support fault termination.

        For details about metrics supported by each instance, see Cloud Product Metrics.

      6. Click Finish.

  7. If a preset scenario in the template is not required, click Delete next to the task. This step is optional.
  8. Click OK.

    After a drill task is created, you can view the drill task and start the drill on the Resilience Center > Chaos Drills > Drill Tasks page.

Starting a Drill Task

Start a drill task.

  1. Log in to COC.
  2. In the navigation pane, choose Resilience Center > Chaos Drills.
  3. Click Drill Tasks.
  4. Locate the drill task you want to start and click Start in the Operation column.
  5. Click OK.

    On the drill details page, you can view the attack progress, including probe installation, drill execution, and environment clearance. The system automatically performs the three steps. The execution time depends on the attack time of the disruptor.

    For probe installation, a probe will be installed on the target server. The probe runs in the system to receive disruptor commands for attack, query, and clearance. For environment clearance, all operations in the system are stopped and removed after the drill is complete or terminated.

  6. Perform the following operations on a drill execution ticket.

    • Terminate: During a drill, click Terminate to stop the task to be executed or the task that is abnormal.
    • Forcible termination: You are advised to use the termination function first. If the termination function fails, you can use the forcible termination function after 5 to 10 minutes. Note that the forcible termination function only closes the current drill ticket and does not automatically clear data in the environment. You need to manually clear data in the environment. For details, see Manually Clearing Data.
    • Retry: If some or all attack tasks fail to check instances, install probes, clear environments, or perform steady-state detection, or if the drill times out, expand the failed attack task and click Retry to retry the task.
    • Skip: If some or all attack tasks fail to be executed during the drill, expand a failed attack task and click Skip to skip the task and execute the next task.
    • View details: Expand an attack task and click Details to view the attack details.

Creating a Drill Report

After a drill task is complete, you can directly create a report if necessary. After a report is created, you can export the report as a PDF file and send it to related personnel. The entire process is flexible and efficient, meeting the requirements of the entire process from production to format fixing.

You can modify the actual fault rectification duration, create improvement tickets, and view fault-handling records in a drill report so that you can comprehensively record and manage drill activities and results.

  1. Log in to COC.
  2. In the navigation pane, choose Resilience Center > Chaos Drills.
  3. Click Drill Tasks.
  4. Locate the target drill task and click Drill Record in the Operation column.
  5. Locate the drill record to be operated and click Generate Report in the Operation column.

    The drill report consists of the following modules:

    • Recovery capability scoring module: You can modify the actual recovery duration. The system automatically generates a recovery capability score.
    • Basic information module: displays basic information about a drill task, including the drill task name, drill report ID, drill start time, end time, drill executor, drill duration, and expected fault recovery duration (minutes).
    • Drill process module: displays drill task cards.
    • Attack task group module: displays attack task details, including attack targets, steady-state metrics, and monitoring metrics.

      The list of selected instances is displayed for attack targets. Line charts are displayed for steady-state metrics and monitoring metrics. If no data is available, No data available is displayed.

    • Improvement item module: You can create improvement tickets. The improvement details are displayed by default, including the processing and verification information.

  6. Click Edit Duration to change the actual recovery duration.

    Actual recovery duration: indicates the actual time required for an application to automatically recover to the normal state or achieve recovery through the execution of a contingency plan after a fault is injected.

    Figure 2 Modifying the actual recovery duration
    Table 10 Parameters for modifying the actual recovery duration

    Parameter

    Description

    Fault Detection Duration (Minutes)

    Enter the fault detection duration.

    Duration from the time when the fault injection is complete until the time when the fault alarm is received.

    Fault Demarcation Duration (Minutes)

    Enter the fault demarcation duration.

    Duration from the time when an alarm is reported until the time when the fault demarcation is complete

    Fault Rectification Duration (Minutes)

    Enter the fault rectification duration.

    Time from fault demarcation until fault rectification.

  7. Click OK.

    After the actual recovery duration is changed, the system automatically generates a recovery capability score.

  8. Click Create Improvement Ticket. In the displayed dialog box, set the improvement ticket information.

    Figure 3 Creating an improvement ticket
    Table 11 Parameters for creating an improvement ticket

    Parameter

    Description

    Improvement Ticket

    Name of an improvement ticket.

    The name can contain a maximum of 64 characters, including letters, digits, hyphens (-), underscores (_), and spaces. It cannot start or end with a space.

    Application

    Select an application for which the improvement is performed.

    Type

    Select an improvement type. You can choose to improve products, O&M, management, or monitoring & alarming.

    Improvement Owner

    Select an owner.

    Improvement Acceptor

    Select an acceptance user.

    Expected Completion Time

    Enter the expected completion time (accurate to day). The selected date cannot be earlier than today.

    Symptom

    Enter the incident-related symptom.

    The value can contain a maximum of 1,000 characters.

    Improvement Ticket Closure Criteria

    Enter the improvement ticket closure criteria.

    The value can contain a maximum of 1,000 characters.

  9. Click OK.

    After the ticket is created, the improvement item list is expanded by default, including the processing and verification information.

Viewing Drill Records

You can view the records of a drill task. A drill task that has not been executed does not contain a record.

  1. Log in to COC.
  2. In the navigation pane, choose Resilience Center > Chaos Drills.
  3. Click Drill Tasks.
  4. Locate the target drill task and click Drill Record in the Operation column.

    The basic information about the drill task includes the drill task name, drill task ID, attack details, and failure mode. All drill records include the drill record ID, execution status, executor, drill start time, and drill end time.

  5. Locate the drill record to be viewed and click View Progress in the Operation column.

    View the attack progress, attack details, and monitoring details of the current drill task.

    • The drill record module displays attack task details, including the progress, task information, and execution time.
    • The attack details module displays the attack status of instances in the application of the current task. BMSs, FlexusL (HCSS) instances, and CSS instances are not supported.
    • The monitoring details module displays real-time monitoring data of attack targets. You need to configure a drill monitoring task when creating an attack task. The screen can be zoomed in to the horizontal full-screen mode.

Manually Clearing Data

Forcibly terminating a drill task only closes the current drill ticket and does not automatically clear data. You need to manually clear the data.

The clearance methods vary from probe to probe. For details, see Table 12.

Table 12 Manually clearing data

Probe Type

Disruptor Type

How to Clear

CFE

All

1. Log in to the console.

2. Go to the /usr/local/cdr_probe and /usr/local/COC-CDR-Probe directories.

3. Run rm -rf /usr/local/cdr_probe/* and rm -rf /usr/local/COC-CDR-Probe/* to delete all files in the directories.

CSS

All

No residual data exists.

DCS

DCS_REDIS_AZSHUTDOWN (powering off a DCS AZ)

Start the DCS instance by referring to DCS Guide.

Other

No residual data exists.

DDS

All

No residual data exists.

Platform

All

No residual data exists.

RDS

RDS_FAILOVER (switching over RDS primary and standby nodes)

No residual data exists.

RDS_SHUTDOWN (stopping an RDS instance)

Start the DCS instance by referring to RDS Guide.

Script

SCRIPT_FAULT (customizing scripts)

1. Cancel the ticket that is not closed (choose Task Management > Execution Records > Script Tickets).

2. Execute the clean method in the custom script.

CCE

All

Delete the namespace and related components from the Kubernetes cluster.

  • Delete the ClusterRoleBinding whose name is cdrprobe.
  • Delete the ClusterRole whose name is cdrprobe.
  • Delete the OperatorCrd whose name is cdrprobes.cdrprobe.io.
  • Delete the Namespace whose name is cdrprobe.

Drill Template Description

In this section, you will find a standard drill template library covering multiple scenarios, including 12 core template categories, such as incident response, process walkthrough, and hands-on contingency plan exercises.

All templates are designed based on industry best practices and feature structural completeness and content reusability. Each includes a standard framework—such as drill context, workflow steps, and role assignments—and supports rapid customization of scenario parameters, risk factors, and response procedures to meet specific operational needs. The instructions and error-prone prompts can help you customize your drill tasks efficiently using these templates, achieving the goal of "ready-to-use, quick-to-deploy" drill preparation.

Table 13 Drill template description

Template

Template Description

Tag

Level

Task Group Name

Attack Scenario

Cross-AZ DR

This drill simulates how a DR failover is performed for the target service and its antecedent middleware when an AZ is faulty or the network is abnormal in the DR deployment architecture.

DR

Advanced

Cross-AZ DR

Server disconnection

Powering off a DCS AZ

Initial Chaos Drill

This is essential for beginners to experience the chaos drill process.

Node

Basic

Initial Chaos Drill

Qualifying practice

High System Resource Usage

This drill specifies the system resource usage to test the service performance in high-pressure scenarios. When host resources are insufficient, you can address the problem in advance.

Node

Medium

Disk Stress

Disk usage increase

Memory Stress

Memory usage increase

CPU Stress

CPU usage increase

HPA Configuration in Kubernetes

In the cloud native architecture, auto scaling is an important feature. This drill simulates scale-up after pod resource usage (such as memory) increases in a short period of time and scale-down after resource usage decreases.

Containers and clusters

Advanced

HPA Configuration in Kubernetes

Pod memory usage increase

Data Storage Exception

Generally, service records are stored on the host or middleware where the service is located. Logs are stored on the disk of the host, and data is stored on the middleware such as DDS. This drill simulates the scenario where the ECS disk I/O is high and the primary and secondary switchover is performed.

Services and data

Medium

Data Storage Exception

Disk I/O pressure increase

Forcibly promoting a secondary node to primary

Automatic Pod Recovery and Scheduling

Kubernetes schedules workloads based on pods. When workloads are generated, the scheduler automatically allocates pods in the workloads. For example, the scheduler distributes pods to nodes that have enough resources.

Cluster

Medium

Automatic Pod Recovery and Scheduling

Memory usage increase

Forcibly stopping a pod

Network Instability Affecting Service Performance

This drill injects a network delay to the NIC of the service host to simulate the impact on services when the network is unstable.

Network

Medium

Network Instability Affecting Service Performance

Network latency

Environment Overload in the Microservice Architecture

Microservices are the mainstream architecture. The core value of microservices is to shorten the service release period and ensure reliable system operation. However, microservices also bring many challenges, such as how to locate and rectify faults in the microservice architecture. This drill simulates overloaded nodes of multiple microservices for your reference.

DR

Medium

Environment Overload in the Microservice Architecture

CPU usage increase

Connection exhaustion

Process killing

Abnormal Server Power-off

This drill simulates whether services can be recovered with no data loss after a server is powered off. In this drill, you can use the corresponding preset contingency plan to recover services after a node is powered off.

Services and data

Medium

Abnormal Server Power-off

Device shutdown

Data Loss in Service Middleware Cache

In large-scale concurrent data query scenarios where high data query efficiency is required, Redis has become an essential service for internet applications due to its significant speed advantages over traditional databases. However, it may face issues related to data consistency and reliability. This chaos drill aims to verify whether service operations remain normal after clearing Redis data.

DR

Medium

Data Loss in Service Middleware Cache

DCS instance restart

Misoperations in the Host Configuration File

It is a high risk for O&M personnel to directly perform black screen operations on the service host. If the permission of the service configuration file is directly modified, the service process may not be able to read or write the file. This chaos drill uses a custom script to perform operations (modifying or removing permissions) on the host configuration file. You can use the prepared contingency plan to recover the service.

Services and data

Medium

Misoperations in the Host Configuration File

Customizing a script

Automatic Workload Switchover

FlexusL instances are next-generation, out-of-the-box, lightweight cloud servers tailored for developers and small- to medium-sized enterprises. You can deploy databases or service applications on FlexusL instances. This drill simulates service workload switchover when processes disappear and database nodes are disconnected.

Network

Advanced

Automatic Workload Switchover

Process killing

Network disconnection

More Operations

After creating a drill task, you can perform the following operations as required.

Table 14 More operations

Function

Scenario

Operation

Modifying a Drill Task

You can modify a created drill task.

Note: If the task has been started and a drill record has been generated, the task cannot be modified.

On the drill task page, locate the target drill task and choose More > Modify in the Operation column.

Deleting a Drill Task

If a created drill task is no longer needed, you can delete it. Do not delete the drill task in the following scenarios:

  • The drill task has generated drill records.
  • The drill task is associated with a drill plan.

On the drill task page, locate the target drill task and choose More > Delete in the Operation column.

Exporting a Report

On the drill report page, click Export Report to download the PDF file.

On the drill report page, click Export Report to download the PDF file.

Refreshing the Drill Reports Page

On the Drill Reports page, click Refresh in the upper right corner to refresh the current page.

On the Drill Reports page, click Refresh in the upper right corner to refresh the current page.

Viewing a Drill Report

On the Drill Records page, locate the row that contains the target drill record and click View Report in the Operation column.

On the Drill Records page, locate the row that contains the target drill record and click View Report in the Operation column.