Node Self-Healing
Node self-healing is an automatic node fault recovery capability provided by CCE. When a specified fault type occurs on a node and persists longer than a defined threshold, CCE automatically triggers a self-healing task that attempts to rectify the fault by restarting the node. This reduces manual intervention and improves cluster availability.
Prerequisites
Advanced self-healing requires the following:
- A Kubernetes cluster meeting the following version requirements is available:
- v1.30: v1.30.14-r100 or later
- v1.31: v1.31.14-r60 or later
- v1.32: v1.32.13-r30 or later
- v1.33: v1.33.12-r10 or later
- v1.34: v1.34.8-r10 or later
- v1.35: v1.35.5-r10 or later
- v1.36: v1.36.2-r0 or later
- Clusters of later versions
- CCE Node Problem Detector has been installed.
Precautions
- Restarting a node evicts or interrupts all pods running on it. Ensure your workloads support HA deployment.
- If the fault persists after the self-healing task is executed, CCE does not retry indefinitely. You must manually locate and resolve the fault.
- If a node fails to self-heal, the node pool will not trigger another self-healing node restart until the fault is rectified.
- In a node pool, only one node can be restarted for self-healing at a time.
- After a node is restarted for self-healing, CCE will not trigger self-healing for that node again within 2 hours.
- Bare metal nodes and HyperNodes do not support node self-healing.
How It Works
Self-healing is implemented by a daemon process and the systemd service manager. The overall process is as follows:
- Scheduled health check: The daemon process runs on the backend and periodically performs health checks on target components. Detection methods include process liveness detection and port response detection.
- Exception determination: If a component fails to respond or returns an error state for multiple consecutive checks, the daemon process flags it as abnormal. A single detection failure does not trigger self-healing, preventing false restarts caused by transient issues such as network jitter.
- Triggering self-healing: Once a component is flagged as abnormal, the daemon process invokes systemd to restart the service, enabling automatic recovery.
If Allow node restart when system and Kubernetes components are abnormal is enabled, CCE enters the node-level self-healing process if the fault persists after the component is restarted.
Self-Healing Process
After a self-healing task is triggered, CCE performs the following steps:
- Marks the affected node as rectifying: The node status is set to Rectifying, and the scheduler stops scheduling new pods to the node.
- Evicts pods on the node: An eviction taint is added to the node to trigger graceful eviction of existing pods. Pods configured to tolerate the taint are not affected. The timeout period is the greater of 10 minutes and the maximum TerminationGracePeriodSeconds of all pods to be evicted on the faulty node. The maximum timeout period is 30 minutes. If eviction fails after the timeout expires, subsequent operations continue regardless.
- Restarts the node: CCE restarts the node to attempt to restore the faulty component.
- Checks fault rectification: After the restart completes, CCE checks whether the fault is rectified.
- Restoration successful: The eviction taint is removed and the node status is restored to normal to accept pod scheduling again.
- Restoration failed: The node remains abnormal, and the self-healing task is marked as failed. Manual intervention is required.
Fault Types
| Fault Type | Description | Threshold | Self-Healing Action |
|---|---|---|---|
| CNI error | The CNI is malfunctioning, affecting pod creation. | 50s | Restart the CNI. |
| kubelet error | The kubelet is unavailable or malfunctioning. The node status cannot be reported. | 50s | Restart kubelet. |
| CRI error | The CRI is malfunctioning. Containers cannot be created or managed. | 50s | Restart containerd or Docker. |
| kube-proxy error | kube-proxy is malfunctioning. | 50s | Restart kube-proxy. |
| PodIdentityAgent error | The PodIdentityAgent is malfunctioning, which may affect CCE add-ons and services that depend on Pod Identity. | 50s | Restart pod-identity-agent. |
| Fault Type | Description | Threshold | Self-Healing Action |
|---|---|---|---|
| CNI error | The CNI is malfunctioning, affecting pod creation. | 50s | Restart the CNI. |
| kubelet error | The kubelet is unavailable or malfunctioning. The node status cannot be reported. | 180s | 1. Restart kubelet (triggered every 50s). 2. If Allow node restart when system and Kubernetes components are abnormal is enabled, restart the affected ECS. |
| CRI error | The CRI is malfunctioning. Containers cannot be created or managed. | 180s | 1. Restart containerd or Docker (triggered every 50s). 2. If Allow node restart when system and Kubernetes components are abnormal is enabled, restart the affected ECS. |
| kube-proxy error | kube-proxy is malfunctioning. | 50s | Restart kube-proxy. |
| PodIdentityAgent error | The PodIdentityAgent is malfunctioning, which may affect CCE add-ons and services that depend on Pod Identity. | 50s | Restart pod-identity-agent. |
| Filesystem read-only (CCE Node Problem Detector v1.19.75 or later) | The node filesystem is read-only and cannot be written. As a result, the kubelet and container runtime cannot work properly. | 180s | If Allow node restart when system and Kubernetes components are abnormal is enabled, restart the affected ECS. |
| systemd offline (CCE Node Problem Detector v1.19.75 or later) | systemd is malfunctioning. Containers cannot be started or destroyed. | 180s | If Allow node restart when system and Kubernetes components are abnormal is enabled, restart the affected ECS. |
| kubeletNotReady (PLEG) | The PLEG health check fails. | 180s | If Allow node restart when system and Kubernetes components are abnormal is enabled, restart the affected ECS. |
Helpful Links
- For details about the constraints, installation, and components of CCE Node Problem Detector, see CCE Node Problem Detector.
Common Issues
What is your overall rating for this page?
Thank you very much for your feedback. We will continue working to improve the documentation.See the reply and handling status in My Cloud VOC.
For any further questions, feel free to contact us through the chatbot.
Chatbot