Updated on 2026-09-03 GMT+08:00

Authorizing O&M on the Event Center Page

Scenario

Huawei Cloud performs fault detection and routine O&M on the software and hardware of nodes. When a node requires hardware maintenance due to an unrecoverable fault, a planned event is pushed to the event center on the console. In the event center, you can view the event information, type, status, and description. You can also authorize Huawei technical support to perform O&M on the faulty node or redeploy the node.

Table 1 Event operation execution conditions

Event Type

Event Status

Operation

Applicable Resource Type

Description

System maintenance

inquiring

Machine maintenance and redeployment

Snt9b, Snt9b21

Machine maintenance requires the host to be powered off. The node will be unavailable during this period. Before maintaining a machine, ensure that services deployed on the node are offline or that stopping the node does not affect services.

Redeployment will cause all data loss on the ECS local disks. If you do not need to retain data on the local disks, you can redeploy the node to quickly restore it.

Local disk recovery

inquiring

Local disk replacement and redeployment

Snt9b, Snt9b21

Replacing a local disk may cause partial or complete data loss on the faulty local disk, and the data cannot be restored. You are advised to back up important data immediately after receiving the event notification.

After the local disk is recovered, you can restore the partition by resetting the node.

During the repair of local disks, you can redeploy the instance to quickly recover services. However, this method will cause all local disk data to be lost and the node to be restarted. You need to back up all local disk data before this operation.

WARNING:

Replacing the local disk will cause local disk data loss. Therefore, migrate services and back up data before authorization.

O&M authorization

inquiring

Node O&M

Snt9b, Snt9b21

Node O&M is to authorize Huawei technical support to perform O&M operations on nodes of the Lite resource type.

Restarting a node

inquiring

Restarting a node

Snt9b, Snt9b21, Snt9b23

A node restart will cause the instance to be stopped, services to be interrupted, and unsaved data to be lost. Confirm the impacts in advance.

Supernode maintenance

inquiring

Machine maintenance

Snt9b23

Machine maintenance requires the host to be powered off. The node will be unavailable during this period. Before maintaining a machine, ensure that services deployed on the node are offline or that stopping the node does not affect services.

Supernode redeployment

inquiring

Redeployment

Snt9b23

Redeployment takes about 10 to 30 minutes and will restart the node. Select a proper time for authorization and switch service traffic in advance.

  • Redeployment will not affect the instance's system disk and EVS data disks.
  • If redeployment is performed forcibly, the system disk and local data disks of the instance will be lost.
  • For instances using local disks, all data stored on the local disks will be deleted after the instance redeployment. To ensure data security, back up local disk data before authorizing the redeployment.

Supernode local disk recovery

inquiring

Replacing a local disk

Snt9b23

When the system detects that data cannot be read from or written to local disks of supernodes due to hardware faults or data exceptions, it will automatically generate a planned event for supernode local disk recovery for the affected nodes.

After the local disk is recovered, you can restore the partition by resetting the node.

Replacing a local disk may cause partial or complete data loss on the faulty local disk, and the data cannot be restored. You are advised to back up important data immediately after receiving the event notification.

WARNING:

Replacing the local disk will cause local disk data loss. Therefore, migrate services and back up data before authorization.

Concepts

  • System maintenance: System maintenance requires the host to be powered off. O&M personnel need to remove and install devices or replace the standby units in the equipment room. Because staff may need to wait for spare parts, maintenance timeliness is limited.
  • Redeployment: Redeployment involves authorizing Huawei technical support to replace the faulty node while ensuring that its configuration is preserved. While this process enables fast fault recovery, note that data stored on the local disk will be lost during the redeployment. Exercise caution when you perform this operation. Migrate services and back up data before redeployment.
  • Silent repair: Automatically handles redeployment failures without manual intervention.

  • Forcible redeployment: Forcibly replaces a node when it is unavailable. All stored data must be cleared.
  • Cold standby node replacement: Replaces a faulty node with a standby node while retaining the configuration information.
  • Local disk: A local disk refers to a physical disk installed on a server. After redeployment, the disk installed on the server will be replaced, resulting in data loss.
  • EVS disk: EVS is a scalable virtual block storage service designed to deliver high reliability and high performance for Elastic Cloud Servers (ECSs) and Bare Metal Servers (BMSs). It offers a wide range of specifications to meet diverse workload needs. EVS disks are not permanently bound to a single server.
  • System disk: A system disk is the primary storage volume that holds the operating system's core components, including the Linux kernel, boot loader, and essential system configuration files. Most system disks are EVS disks. Currently, the system disk is typically implemented as an EVS disk.

Constraints

  • Currently, all regions on the Huawei Cloud Chinese Mainland website and some regions (CN-Hong Kong, AF-Johannesburg, and AP-Singapore) on the Huawei Cloud International website support this function.
  • Only Snt9b nodes, Snt9b21 supernodes, and Snt9b23 supernodes support hardware maintenance through planned events.
  • Redeployment of supernodes must be performed within physical supernodes. If a supernode is in its full configuration (48 nodes), redeployment is not supported. Instead, a planned event for supernode system maintenance is directly pushed.
  • If the planned event does not meet the event status requirements listed in Table 1, the authorization button is unavailable.
  • After the local node disk and supernode disk are restored, the local disk data will be lost. Therefore, migrate services and backup data before authorization. After the local disk is restored, log in to the Lite Server to partition the local disk.

Status Transition of a Planned Event

After the planned event is authorized, its status changes from Authorization Pending to Execution Pending, Executing, and then Completed or Failed in sequence.

In the planned event list, events are retained based on the event status.

  • If the status is Canceled or Completed, the events are retained for seven days.
  • If the status is Authorization Pending, events are retained for 30 days. If the events are not authorized within 30 days, the status is automatically changed to Canceled.
  • If the status is Failed, contact Huawei technical support.
Table 2 Planned event statuses

Status

Description

Supported Operations

Authorization Pending

The planned event has been generated but has not been authorized. Authorization is required.

Authorize the planned event and select the corresponding handling method.

Execution Pending

The planned event has been authorized and will be automatically executed by the system based on the execution window.

No operation is required.

Executing

The planned event is being executed according to the handling method selected by the user.

No operation is required.

Completed

The planned event has been successfully executed, and the corresponding fault has been rectified.

No operation is required.

Canceled

The planned event is canceled by the system. For example, if the root cause fault that triggers the planned event has been rectified, the system automatically cancels the planned event.

No operation is required.

Failed

The execution of the planned event failed.

Contact Huawei technical support for O&M.

  • If you select redeployment, the status of the planned event will change from inquiring to executing because the planned event is immediately executed after authorization.
  • Forcible redeployment resets the node, deleting all data on both its local and system disks. Exercise caution when performing this operation.

Procedure

If the faulty nodes meet the requirements listed in Table 1, you can authorize Huawei technical support to perform O&M on the faulty nodes.
  1. Log in to the ModelArts console. In the navigation pane on the left, choose Resource Management > Auxiliary Tools > Event Center.
  2. Click Authorize Handling in the Operation column. In the displayed dialog box, select a handling method and click OK.
    Figure 1 Authorize handling

  3. After the authorization, Authorize Handling becomes unavailable, and the event status changes to Execution Pending or Executing.

    If redeployment is selected, the progress is displayed in the Operation column. Click the View Progress button in the Operation column to view the redeployment progress. If the redeployment fails, you can view the failure cause. If you select Silent repair during redeployment, you do not need to pay attention to it. The system automatically switches to the in-place repair planned event or redeploys the node.

After the O&M operation is complete, the event status changes to Completed. No further action is required.

Operation Examples in System Maintenance Scenarios

  • Snt9b23 supernode system maintenance

    Machine maintenance is to authorize Huawei O&M personnel to repair hardware devices. This process requires equipment room personnel to replace spare parts, which takes a certain period of time. During system maintenance, the node may restart or power off, making it temporarily unavailable. Before authorizing, ensure that your workloads have been migrated and local data has been backed up.

    If you are sure that your workloads will not be affected, you can enter YES in the authorization confirmation dialog box, or click Auto Enter and then click OK to complete the authorization process.

  • Snt9b node and Snt9b21 supernode system maintenance

    For system maintenance events of Snt9b nodes and Snt9b21 supernodes, you can choose to maintain the machine or redeploy the node. The following figure uses system maintenance as an example.

    Figure 2 Authorize Handling

    The handling method is the same as that for Snt9b23. If you are sure that your workloads will not be affected, you can enter YES in the authorization confirmation dialog box, or click Auto Enter and then click OK to complete the authorization process.

Operation Examples in Redeployment Scenarios

Dedicated resource pools and Lite Clusters use the same processing logic during redeployment. However, Lite Servers support only redeployment in silent repair mode.

  • If redeployment is selected without silent repair, when a task fails, for example, no cold standby server of the same specifications is available, contact Huawei technical support.
  • Dedicated resource pool and Lite Cluster

    Example 1: Snt9b23 supernode redeployment

    1. Procedure
      • Click Authorize Handling in 2 and select Redeployment. The system automatically replaces the cold standby server. The node IP address, configuration, and source node are exactly the same as before. The restoration takes about 10 to 30 minutes.
      • Preparations
        • Set Silent repair and Forcible redeployment, enter YES, and click OK.
    2. Silent repair mechanism
      • Enabled: If redeployment fails (for example, no cold standby server is available or the admission check fails), the system automatically transfers the task to the system maintenance phase or initiates redeployment for multiple times. A new plan event is generated without the need to authorize again.
      • Disabled: If the planned event fails, contact technical support.
    3. Forcible redeployment description
      • Application: Forcibly perform cold standby replacement when the node is unavailable, or the user needs to reset the node.
      • Notes:
        • The operation takes 20 to 30 minutes. Data on the local disk and system disk will be cleared. Back up the data in advance.
        • If the event status is Completed, the fault is rectified and the node can be scheduled for service jobs. If the event status is Failed, contact technical support.
    4. Silent repair association logic

      If a redeployment event is in the Canceled state on the same node, the system automatically generates a supernode maintenance event. Secondary authorization is not required.

    Example 2: Snt9b and Snt9b21 supernode redeployment

    1. Dependency

      The operation depends on system maintenance or local disk recovery events.

    2. Procedure
      • Click Authorize Handling in 2 and select Redeployment. Enable both Silent repair and Forcible redeployment.
      • Under Confirm Authorization, enter YES or click Auto Enter, and click OK.
  • Lite Server

    Lite Servers support silent repair. You only need to authorize redeployment by referring to Procedure.

    After the redeployment is complete, reset the node.

FAQ

  1. How do I locate a faulty node in a dedicated resource pool?

    In a dedicated resource pool, ModelArts will add a taint to a faulty Kubernetes node so that jobs will not be scheduled to the tainted node. You can locate the fault by referring to the isolation code and detection method. For details, see Faulty Nodes in a Resource Pool.

  2. What can I do if the status on the node details page is inconsistent with the actual status when a Lite Server is being redeployed?

    Lite Server status is not synchronized in real time. The automatic synchronization period is long. As a result, the status on the node details page may not be the actual status. To address this issue, click Synchronize in the upper right corner of the Lite Server details page.

    Example:

    Snt9b23 automatically shuts down before redeployment. After the shutdown, if the redeployment fails, for example, there is no standby server, the actual node is shut down. However, the Lite Server may still be displayed as Running before automatic synchronization. You can manually synchronize the status.