Authorizing O&M on the Event Center Page
Scenario
Huawei Cloud performs fault detection and routine O&M on the software and hardware of nodes. When a node requires hardware maintenance due to an unrecoverable fault, a planned event is pushed to the event center on the console. In the event center, you can view the event information, type, status, and description. You can also authorize Huawei technical support to perform O&M on the faulty node or redeploy the node.
| Event Type | Event Status | Operation | Applicable Resource Type | Description |
|---|---|---|---|---|
| System maintenance | inquiring | Machine maintenance and redeployment | Snt9b, Snt9b21 | Machine maintenance requires the host to be powered off. The node will be unavailable during this period. Before maintaining a machine, ensure that services deployed on the node are offline or that stopping the node does not affect services. Redeployment will cause all data loss on the ECS local disks. If you do not need to retain data on the local disks, you can redeploy the node to quickly restore it. |
| Local disk recovery | inquiring | Local disk replacement and redeployment | Snt9b, Snt9b21 | Replacing a local disk may cause partial or complete data loss on the faulty local disk, and the data cannot be restored. You are advised to back up important data immediately after receiving the event notification. After the local disk is recovered, you can restore the partition by resetting the node. During the repair of local disks, you can redeploy the instance to quickly recover services. However, this method will cause all local disk data to be lost and the node to be restarted. You need to back up all local disk data before this operation. WARNING: Replacing the local disk will cause local disk data loss. Therefore, migrate services and back up data before authorization. |
| O&M authorization | inquiring | Node O&M | Snt9b, Snt9b21 | Node O&M is to authorize Huawei technical support to perform O&M operations on nodes of the Lite resource type. |
| Restarting a node | inquiring | Restarting a node | Snt9b, Snt9b21, Snt9b23 | A node restart will cause the instance to be stopped, services to be interrupted, and unsaved data to be lost. Confirm the impacts in advance. |
| Supernode maintenance | inquiring | Machine maintenance | Snt9b23 | Machine maintenance requires the host to be powered off. The node will be unavailable during this period. Before maintaining a machine, ensure that services deployed on the node are offline or that stopping the node does not affect services. |
| Supernode redeployment | inquiring | Redeployment | Snt9b23 | Redeployment takes about 10 to 30 minutes and will restart the node. Select a proper time for authorization and switch service traffic in advance.
|
| Supernode local disk recovery | inquiring | Replacing a local disk | Snt9b23 | When the system detects that data cannot be read from or written to local disks of supernodes due to hardware faults or data exceptions, it will automatically generate a planned event for supernode local disk recovery for the affected nodes. After the local disk is recovered, you can restore the partition by resetting the node. Replacing a local disk may cause partial or complete data loss on the faulty local disk, and the data cannot be restored. You are advised to back up important data immediately after receiving the event notification. WARNING: Replacing the local disk will cause local disk data loss. Therefore, migrate services and back up data before authorization. |
Concepts
- System maintenance: System maintenance requires the host to be powered off. O&M personnel need to remove and install devices or replace the standby units in the equipment room. Because staff may need to wait for spare parts, maintenance timeliness is limited.
- Redeployment: Redeployment involves authorizing Huawei technical support to replace the faulty node while ensuring that its configuration is preserved. While this process enables fast fault recovery, note that data stored on the local disk will be lost during the redeployment. Exercise caution when you perform this operation. Migrate services and back up data before redeployment.
-
Silent repair: Automatically handles redeployment failures without manual intervention.
- Forcible redeployment: Forcibly replaces a node when it is unavailable. All stored data must be cleared.
- Cold standby node replacement: Replaces a faulty node with a standby node while retaining the configuration information.
- Local disk: A local disk refers to a physical disk installed on a server. After redeployment, the disk installed on the server will be replaced, resulting in data loss.
- EVS disk: EVS is a scalable virtual block storage service designed to deliver high reliability and high performance for Elastic Cloud Servers (ECSs) and Bare Metal Servers (BMSs). It offers a wide range of specifications to meet diverse workload needs. EVS disks are not permanently bound to a single server.
- System disk: A system disk is the primary storage volume that holds the operating system's core components, including the Linux kernel, boot loader, and essential system configuration files. Most system disks are EVS disks. Currently, the system disk is typically implemented as an EVS disk.
Constraints
- Currently, all regions on the Huawei Cloud Chinese Mainland website and some regions (CN-Hong Kong, AF-Johannesburg, and AP-Singapore) on the Huawei Cloud International website support this function.
- Only Snt9b nodes, Snt9b21 supernodes, and Snt9b23 supernodes support hardware maintenance through planned events.
- Redeployment of supernodes must be performed within physical supernodes. If a supernode is in its full configuration (48 nodes), redeployment is not supported. Instead, a planned event for supernode system maintenance is directly pushed.
- If the planned event does not meet the event status requirements listed in Table 1, the authorization button is unavailable.
- After the local node disk and supernode disk are restored, the local disk data will be lost. Therefore, migrate services and backup data before authorization. After the local disk is restored, log in to the Lite Server to partition the local disk.
Status Transition of a Planned Event
After the planned event is authorized, its status changes from Authorization Pending to Execution Pending, Executing, and then Completed or Failed in sequence.
In the planned event list, events are retained based on the event status.
- If the status is Canceled or Completed, the events are retained for seven days.
- If the status is Authorization Pending, events are retained for 30 days. If the events are not authorized within 30 days, the status is automatically changed to Canceled.
- If the status is Failed, contact Huawei technical support.
| Status | Description | Supported Operations |
|---|---|---|
| Authorization Pending | The planned event has been generated but has not been authorized. Authorization is required. | Authorize the planned event and select the corresponding handling method. |
| Execution Pending | The planned event has been authorized and will be automatically executed by the system based on the execution window. | No operation is required. |
| Executing | The planned event is being executed according to the handling method selected by the user. | No operation is required. |
| Completed | The planned event has been successfully executed, and the corresponding fault has been rectified. | No operation is required. |
| Canceled | The planned event is canceled by the system. For example, if the root cause fault that triggers the planned event has been rectified, the system automatically cancels the planned event. | No operation is required. |
| Failed | The execution of the planned event failed. | Contact Huawei technical support for O&M. |
- If you select redeployment, the status of the planned event will change from inquiring to executing because the planned event is immediately executed after authorization.
- Forcible redeployment resets the node, deleting all data on both its local and system disks. Exercise caution when performing this operation.
Procedure
- Log in to the ModelArts console. In the navigation pane on the left, choose Resource Management > Auxiliary Tools > Event Center.
- Click Authorize Handling in the Operation column. In the displayed dialog box, select a handling method and click OK. Figure 1 Authorize handling
- After the authorization, Authorize Handling becomes unavailable, and the event status changes to Execution Pending or Executing.
If redeployment is selected, the progress is displayed in the Operation column. Click the View Progress button in the Operation column to view the redeployment progress. If the redeployment fails, you can view the failure cause. If you select Silent repair during redeployment, you do not need to pay attention to it. The system automatically switches to the in-place repair planned event or redeploys the node.
After the O&M operation is complete, the event status changes to Completed. No further action is required.
Operation Examples in System Maintenance Scenarios
- Snt9b23 supernode system maintenance
Machine maintenance is to authorize Huawei O&M personnel to repair hardware devices. This process requires equipment room personnel to replace spare parts, which takes a certain period of time. During system maintenance, the node may restart or power off, making it temporarily unavailable. Before authorizing, ensure that your workloads have been migrated and local data has been backed up.
If you are sure that your workloads will not be affected, you can enter YES in the authorization confirmation dialog box, or click Auto Enter and then click OK to complete the authorization process.
- Snt9b node and Snt9b21 supernode system maintenance
For system maintenance events of Snt9b nodes and Snt9b21 supernodes, you can choose to maintain the machine or redeploy the node. The following figure uses system maintenance as an example.
Figure 2 Authorize Handling
The handling method is the same as that for Snt9b23. If you are sure that your workloads will not be affected, you can enter YES in the authorization confirmation dialog box, or click Auto Enter and then click OK to complete the authorization process.
Operation Examples in Redeployment Scenarios
Dedicated resource pools and Lite Clusters use the same processing logic during redeployment. However, Lite Servers support only redeployment in silent repair mode.
- If redeployment is selected without silent repair, when a task fails, for example, no cold standby server of the same specifications is available, contact Huawei technical support.
- Dedicated resource pool and Lite Cluster
Example 1: Snt9b23 supernode redeployment
- Procedure
- Click Authorize Handling in 2 and select Redeployment. The system automatically replaces the cold standby server. The node IP address, configuration, and source node are exactly the same as before. The restoration takes about 10 to 30 minutes.
- Preparations
- Set Silent repair and Forcible redeployment, enter YES, and click OK.
- Silent repair mechanism
- Enabled: If redeployment fails (for example, no cold standby server is available or the admission check fails), the system automatically transfers the task to the system maintenance phase or initiates redeployment for multiple times. A new plan event is generated without the need to authorize again.
- Disabled: If the planned event fails, contact technical support.
- Forcible redeployment description
- Application: Forcibly perform cold standby replacement when the node is unavailable, or the user needs to reset the node.
- Notes:
- The operation takes 20 to 30 minutes. Data on the local disk and system disk will be cleared. Back up the data in advance.
- If the event status is Completed, the fault is rectified and the node can be scheduled for service jobs. If the event status is Failed, contact technical support.
- Silent repair association logic
If a redeployment event is in the Canceled state on the same node, the system automatically generates a supernode maintenance event. Secondary authorization is not required.
Example 2: Snt9b and Snt9b21 supernode redeployment
- Dependency
The operation depends on system maintenance or local disk recovery events.
- Procedure
- Click Authorize Handling in 2 and select Redeployment. Enable both Silent repair and Forcible redeployment.
- Under Confirm Authorization, enter YES or click Auto Enter, and click OK.
- Procedure
- Lite Server
Lite Servers support silent repair. You only need to authorize redeployment by referring to Procedure.
After the redeployment is complete, reset the node.
FAQ
- How do I locate a faulty node in a dedicated resource pool?
In a dedicated resource pool, ModelArts will add a taint to a faulty Kubernetes node so that jobs will not be scheduled to the tainted node. You can locate the fault by referring to the isolation code and detection method. For details, see Faulty Nodes in a Resource Pool.
- What can I do if the status on the node details page is inconsistent with the actual status when a Lite Server is being redeployed?
Lite Server status is not synchronized in real time. The automatic synchronization period is long. As a result, the status on the node details page may not be the actual status. To address this issue, click Synchronize in the upper right corner of the Lite Server details page.
Example:
Snt9b23 automatically shuts down before redeployment. After the shutdown, if the redeployment fails, for example, there is no standby server, the actual node is shut down. However, the Lite Server may still be displayed as Running before automatic synchronization. You can manually synchronize the status.
What is your overall rating for this page?
Thank you very much for your feedback. We will continue working to improve the documentation.See the reply and handling status in My Cloud VOC.
For any further questions, feel free to contact us through the chatbot.
Chatbot