CCE AI Suite (Ascend NPU)
Introduction
The CCE AI Suite (Ascend NPU) add-on supports and manages Huawei NPUs in containers. It provides functions such as automatic driver installation, device registration and scheduling, performance monitoring, and virtual resource management. With this add-on, you can implement automatic deployment, refined scheduling, visualized monitoring, and virtualization of NPUs for diverse heterogeneous computing requirements. It is typically used in:
- AI model training: It supports multi-NPU parallelism and refined resource scheduling to improve the efficiency and stability of large-scale model training.
- AI inference and real-time services: It supports low-latency, on-demand computing to ensure real-time performance and service availability of inference tasks.
- Resource isolation environments: It supports vNPU management to divide resources by granularity for compute isolation and quota control in multi-user scenarios.
-
Training task monitoring and resource optimization: It supports metric collection and visual analysis to monitor training task performance and optimize resource usage.
After this add-on is installed, you can create Ascend-accelerated nodes to quickly and efficiently process inference and image recognition.
A5 HyperNodes support parameter plane RDMA device management. The enable_rdma_shared_dp field in the add-on configuration controls whether to deploy the RDMA Shared Device Plugin. This plugin enables the discovery, allocation, and mounting of RDMA devices on the UB bus, as well as the detection and reporting of NIC faults. For details, see Configuring the RDMA Capability.
Constraints
- To use Ascend-accelerated nodes in a cluster, the CCE AI Suite (Ascend NPU) add-on must be installed.
- After an AI-accelerated node is migrated, the node will be reset. If the automatic driver installation function is enabled (supported only by the add-on of version 1.2.5 or later) and the driver corresponding to the NPU node model is selected for the CCE AI Suite (Ascend NPU) add-on of the destination cluster, the NPU driver will be automatically installed after the node migration.
- If the add-on version is earlier than 1.2.8, you need to manually restart the node after the driver installation for the driver to work.
- If the add-on version is 1.2.8 or later, the node will be automatically restarted after the driver is installed. After the restart, the driver works.
- If the NPU driver installation command is included in the post-installation script of the node pool, with automatic driver installation enabled and the correct NPU driver model selected, driver installation commands from the frontend and the npu-driver-installer pod will be executed concurrently to install drivers during node pool scale-out. This may result in discrepancies in the installed driver or installation failures. Therefore, when the driver selection function is enabled for huawei-npu, you are not advised to scale out a node pool that has a compiled post-installation command or post-installation command when creating a node pool to install the NPU driver.
- Access to Ascend NPUs is controlled by the HwHiAiUser group created during driver installation. The NPU runtime user in the container (whether root or non-root) must belong to this group. Otherwise, npu-smi commands will fail and the NPU cannot be used for inference or training.
- RDMA device management:
- RDMA device management is available only in CCE standard and Turbo clusters.
- The RDMA driver must be installed on any node that uses RDMA functionality, and kernel modules such as ib_uverbs and ib_core must be loaded.
- The /dev/infiniband/ directory must exist on the node and contain RDMA device files.
- The RDMA Shared Device Plugin is deployed as a DaemonSet and requires privileged permissions and hostNetwork.
- RDMA device management supports only Huawei UB bus devices (vendor=0xcc08, deviceID=0x8200).
- To enable RDMA, set enable_rdma_shared_dp to true. By default, RDMA is disabled.
Installing the Add-on
- Log in to the CCE console and click the cluster name to access the cluster console.
- In the navigation pane, choose Add-ons. In the right pane, find the CCE AI Suite (Ascend NPU) add-on and click Install.
- In the Metric-based Observation area, enable Use NPU-Exporter to Observe NPU Metrics. After this function is enabled, NPU-Exporter will be deployed on the NPU nodes as a DaemonSet. When using this component, pay attention to the following points:
- NPU-Exporter is supported when the add-on version is 2.1.55 or later. In addition, it must be used with the NPU driver of version 24.x or later.
- After NPU-Exporter is enabled, if you need to report the collected NPU monitoring data to AOM, see Comprehensive Monitoring of NPU Metrics.
- Determine whether to enable Auto Driver Installation (supported only when the add-on version is 1.2.5 or later).
- Enabled: You can specify the driver version based on the NPU model for easier driver maintenance. You can determine whether to enable this function based on different models. After the driver is enabled, the add-on automatically installs the driver based on the specified driver version. By default, the recommended driver is used. You can also select Path to a custom driver from the drop-down list and enter a driver address.
- The add-on installs the driver based on the driver version selected for the specified model. Such installation is only for nodes with no NPU driver installed. Nodes with an NPU driver installed remain unchanged. If you change the driver version when upgrading the add-on or updating add-on parameters, such change takes effect only on the nodes with no NPU driver installed.
- After the driver is successfully installed, the node automatically restarts. To prevent service loss during the restart, you are advised to drain the node in advance and then install or upgrade the driver. For details about how to drain a node, see Draining a Node. For details about how to verify the driver installation, see How to Check Whether the NPU Driver Has Been Installed on a Node.
- Uninstalling the add-on does not automatically delete the installed NPU driver. For details about how to uninstall the NPU driver, see Uninstalling the NPU Driver.
- Disabled: Driver versions are decided by the system, and the drivers cannot be maintained using the add-on. When you add an NPU node on the console, the system adds the command to install an NPU driver (version and type decided by the system) and automatically restarts the node after the driver installation is complete. Adding an NPU node in another way, such as using an API, requires you to add the driver installation command to Post-installation Script.
- The following table lists what NPUs and OS specifications are supported.
Table 1 Specification adaptation NPU Type
Supported OS
Snt3 (ascend-snt3)
EulerOS 2.5 x86, CentOS 7.6 x86, EulerOS 2.9 x86, and EulerOS 2.8 Arm
NOTE:The Snt3 Arm model supports up to EulerOS 2.8 Arm, which has now reached EOS. For details, see EOS Plan.
CCE standard and Turbo clusters v1.28 and later do not support Arm EulerOS 2.8. To use NPUs in these clusters, select compatible NPUs and OSs by referring to Mappings Between Cluster Versions and OS Versions and Software Versions Required by Different Models. For details about the purchase process, see Using Lite Cluster.
- Enabled: You can specify the driver version based on the NPU model for easier driver maintenance.
- Click Install.
Components
| Component | Description | Resource Type |
|---|---|---|
| npu-driver-installer | Used for installing an NPU driver on NPU nodes. | DaemonSet |
| npu-exporter | Used for monitoring and collecting NPU metric data. You need to manually enable it. After this component is enabled, it runs as a DaemonSet on each NPU node. | DaemonSet |
| ascend-vnpu-manager | Enables node pool–level NPU virtualization and allows CCE to create vNPUs. By doing so, it facilitates more efficient utilization of NPU resources. After installing the add-on, go to Settings, click the Heterogeneous Resources tab, and enable NPU virtualization. The add-on will then deploy the component on the required nodes. For details, see Static NPU Virtualization. | DaemonSet |
NPU Metrics
In CCE standard and Turbo clusters of v1.25 or later, the NPU metrics listed in the table below are exposed through device-plugin. They can be collected, reported, and displayed on AOM. For details about how to collect and monitor more NPU metrics, see Comprehensive Monitoring of NPU Metrics.
| Metric | Monitoring Level | Remarks |
|---|---|---|
| cce_npu_memory_total | NPU cards | Total NPU memory |
| cce_npu_memory_used | NPU cards | NPU memory usage |
| cce_npu_utilization | NPU cards | NPU compute usage |
How to Check Whether the NPU Driver Has Been Installed on a Node
After ensuring that the driver is successfully installed, restart the node for the driver to take effect. Otherwise, the driver cannot take effect and NPU resources are unavailable. To check whether the driver is installed, perform the following operations:
- On the Add-ons page, click CCE AI Suite (Ascend NPU).

- Verify that the node where npu-driver-installer is deployed is in the Running state.

If the node is restarted before the NPU driver is installed, the driver installation may fail, and a message is displayed on the Nodes page indicating that the driver is not ready. In this case, uninstall the NPU driver from the node and restart the npu-driver-installer pod to reinstall the NPU driver. After confirming that the driver is installed, restart the node. For details about how to uninstall the driver, see Uninstalling the NPU Driver.
Uninstalling the NPU Driver
Log in to the node, obtain the driver operation records in the /var/log/ascend_seclog/operation.log file, and find the driver run package used in the last installation. If the log file does not exist, the driver is installed using the npu_x86_latest.run or npu_arm_latest.run driver combined package. After finding the driver installation package, run the bash {run package name} --uninstall command to uninstall the driver and restart the node as prompted.
- Log in to the node where the NPU driver needs to be uninstalled and find the /var/log/ascend_seclog/operation.log file.
- If the /var/log/ascend_seclog/operation.log file can be found, view the driver installation log to find the driver installation record.

If the /var/log/ascend_seclog/operation.log file cannot be found, the driver may be installed using the npu_x86_latest.run or npu_arm_latest.run driver combined package. You can confirm this by checking whether the /usr/local/HiAI/driver/ directory exists.
The combined package of the NPU driver is stored in the /root/d310_driver directory; other driver installation packages are stored in the /root/npu-drivers directory.
- After finding the driver installation package, run the bash {run package path} --uninstall command to uninstall the driver. The following uses Ascend310-hdk-npu-driver_6.0.rc1_linux-x86-64.run as an example:
bash /root/npu-drivers/Ascend310-hdk-npu-driver_6.0.rc1_linux-x86-64.run --uninstall

- Restart the node as prompted. (The installation and uninstallation of the current NPU driver take effect only after the node is restarted.)
Enabling RDMA
| Parameter | Description | Default Value |
|---|---|---|
| enable_rdma_shared_dp | Controls whether to deploy the RDMA Shared Device Plugin. | false |
If enable_rdma_shared_dp is set to true, you must also configure the rdma_shared_dp configuration block, which involves the following parameters.
| Parameter | Description | Default Value |
|---|---|---|
| periodicUpdateInterval | Interval for updating the device status, in seconds. | 300 |
| faultDetectPeriod | Fault detection period, in minutes. | 5 |
| affinity | Node affinity settings | [{"key":"accelerator/huawei-npu","operator":"Exists"}] |
| configList[].resourcePrefix | Kubernetes resource name prefix. | huawei.com |
| configList[].resourceName | Kubernetes resource name. | ub_rdma |
| configList[].rdmaHcaMax | Maximum number of resource instances that can be allocated to a container from a single physical HCA device. | 8 |
| configList[].selectors.buses | Bus type. The value is fixed at ub. | ub |
| configList[].selectors.vendors | Device vendor ID. | 0xcc08 |
| configList[].selectors.deviceIDs | Device ID. | 0x8200 |
The following shows an example configuration:
{
"periodicUpdateInterval": 300,
"faultDetectPeriod": 5,
"configList": [
{
"resourcePrefix": "huawei.com",
"resourceName": "ub_rdma",
"rdmaHcaMax": 8,
"selectors": {
"buses": ["ub"],
"vendors": ["0xcc08"],
"deviceIDs": ["0x8200"]
}
}
]
} Components
| Component | Description | Resource Type |
|---|---|---|
| k8s-rdma-shared-dev-plugin | RDMA shared device management add-on, used to discover, allocate, and mount RDMA devices on the UB bus, and detect faults of the 1825 NIC. The health check port is 11257. | DaemonSet |
Component Functions
The k8s-rdma-shared-dev-plugin component provides the following functions:
- Device discovery: automatically discovers RDMA devices on nodes based on configured selectors such as vendor ID, device ID, and bus type.
- Resource registration: registers RDMA device resources with Kubernetes and supports device sharing mode.
- Device allocation: responds to device allocation requests from kubelet and allocates RDMA devices to containers.
- Periodic update: rescans devices and updates the resource list at the interval specified by periodicUpdateInterval.
- Fault detection: performs fault detection at the interval specified by faultDetectPeriod.
- Fault reporting: reports device fault information to ConfigMap (dpuinfo-{nodeName}) for the Volcano scheduler to use.
Health Check
| Parameter | Value |
|---|---|
| Check Method | HTTP GET |
| Check Port | 11257 |
| Check Path | N/A |
| Initial Delay | 30 seconds |
| Check Interval | 15 seconds |
| Timeout | 5 seconds |
| Failure Threshold | 3 |
If the check fails three consecutive times (at most 65 seconds), kubelet automatically restarts the container.
Volume Mounting
| Container Mount Path | Host Path | Description |
|---|---|---|
| /var/lib/kubelet/device-plugins | /var/lib/kubelet/device-plugins | Communication directory of the kubelet device add-on |
| /k8s-rdma-shared-dev-plugin | ConfigMap rdma-devices-config | RDMA device configuration file |
| /dev/ | /dev/ | Device file directory |
| /sys | /sys/ | sysfs virtual file system, which is used for device enumeration |
| /dev/infiniband | /dev/infiniband/ | InfiniBand device files |
| /var/log/cce/kubernetes | /var/log/cce/kubernetes | CCE component log directory |
| /usr/sbin/hinicadm5 | /usr/sbin/hinicadm5 | 1825 NIC management tool (read-only) |
| /var/log/hinic5 | /var/log/hinic5 | 1825 NIC log directory |
Viewing Logs
The RDMA add-on logs are stored in the /var/log/cce/kubernetes/ directory on the node. The log file name is k8s-rdma-shared-dp.log.
Run the following commands to view logs:
# View the pod logs of the RDMA component. kubectl logs -n kube-system -l k8s-app=rdma-shared-dp-ds # View the log file on the node. cat /var/log/cce/kubernetes/k8s-rdma-shared-dp.log
Using a Huawei RDMA Device in a Pod
apiVersion: v1
kind: Pod
metadata:
name: rdma-test-pod
spec:
hostNetwork: true # The RDMA Shared Device Plugin must be used with the host network.
containers:
- name: rdma-container
image: your-rdma-application:latest
resources:
requests:
huawei.com/ub_rdma: '1'
limits:
huawei.com/ub_rdma: '1'
- hostNetwork: true must be configured. Pods can access RDMA devices only through the host network namespace.
- In the Allocate phase, the add-on automatically mounts necessary RDMA device files (such as /dev/infiniband) into the container. You do not need to manually configure volumeMounts.
Helpful Links
Release History
| Add-on Version | Supported Cluster Version | New Feature |
|---|---|---|
| 2.8.11 | v1.29 v1.30 v1.31 v1.32 v1.33 v1.34 v1.35 v1.36 | Added support for CANN 8.5.0 in FlexNPU. |
| 2.7.0 | v1.29 v1.30 v1.31 v1.32 v1.33 v1.34 v1.35 v1.36 | Supported CCE clusters v1.36. |
| 2.5.1 | v1.29 v1.30 v1.31 v1.32 v1.33 v1.34 v1.35 | Supported CCE clusters v1.35. |
| 2.4.2 | v1.28 v1.29 v1.30 v1.31 v1.32 v1.33 v1.34 | Supported Snt9b23 devices that do not have RDMA over Converged Ethernet (RoCE) network interfaces. |
| 2.4.1 | v1.28 v1.29 v1.30 v1.31 v1.32 v1.33 v1.34 | Supported CCE clusters v1.34. |
| 2.3.2 | v1.27 v1.28 v1.29 v1.30 v1.31 v1.32 v1.33 | Added image signature. |
| 2.2.2 | v1.27 v1.28 v1.29 v1.30 v1.31 v1.32 v1.33 | Supported CCE clusters v1.33. |
| 2.1.63 | v1.25 v1.27 v1.28 v1.29 v1.30 v1.31 v1.32 | Supported CCE clusters v1.32. |
| 2.1.53 | v1.25 v1.27 v1.28 v1.29 v1.30 v1.31 | Fixed the security vulnerabilities. |
| 2.1.46 | v1.21 v1.23 v1.25 v1.27 v1.28 v1.29 v1.30 v1.31 | Supported CCE clusters v1.31. |
| 2.1.23 | v1.21 v1.23 v1.25 v1.27 v1.28 v1.29 v1.30 | Fixed some issues. |
| 2.1.22 | v1.21 v1.23 v1.25 v1.27 v1.28 v1.29 v1.30 |
|
| 2.1.14 | v1.21 v1.23 v1.25 v1.27 v1.28 v1.29 v1.30 | Fixed some issues. |
| 2.1.7 | v1.21 v1.23 v1.25 v1.27 v1.28 v1.29 | Resolved the issue that npu-smi fails to be automatically mounted to a service container. |
| 2.1.5 | v1.21 v1.23 v1.25 v1.27 v1.28 v1.29 |
|
| 2.0.9 | v1.21 v1.23 v1.25 v1.27 v1.28 | Resolved the issue that process-level fault recovery and adding annotations to workloads occasionally failed. |
| 2.0.5 | v1.21 v1.23 v1.25 v1.27 v1.28 |
|
| 1.2.14 | v1.19 v1.21 v1.23 v1.25 v1.27 | Supported NPU monitoring. |
| 1.2.6 | v1.19 v1.21 v1.23 v1.25 | Supported automatic installation of NPU drivers. |
| 1.2.5 | v1.19 v1.21 v1.23 v1.25 | Supported automatic installation of NPU drivers. |
| 1.2.4 | v1.19 v1.21 v1.23 v1.25 | Supported CCE clusters v1.25. |
| 1.2.2 | v1.19 v1.21 v1.23 | Supported CCE clusters v1.23. |
| 1.2.1 | v1.19 v1.21 v1.23 | Supported CCE clusters v1.23. |
| 1.1.8 | v1.15 v1.17 v1.19 v1.21 | Supported CCE clusters v1.21. |
| 1.1.2 | v1.15 v1.17 v1.19 | Added the default seccomp profile. |
| 1.1.1 | v1.15 v1.17 v1.19 | Supported CCE clusters v1.15. |
| 1.1.0 | v1.17 v1.19 | Supported CCE clusters v1.19. |
| 1.0.8 | v1.13 v1.15 v1.17 | Adapted to the Snt3 C75 drivers. |
| 1.0.6 | v1.13 v1.15 v1.17 | Supported the C75 drivers. |
| 1.0.5 | v1.13 v1.15 v1.17 | Allowed containers to use Huawei NPUs. |
| 1.0.3 | v1.13 v1.15 v1.17 | Allowed containers to use Huawei NPUs. |
What is your overall rating for this page?
Thank you very much for your feedback. We will continue working to improve the documentation.See the reply and handling status in My Cloud VOC.
For any further questions, feel free to contact us through the chatbot.
Chatbot