Updated on 2026-09-29 GMT+08:00

CCE AI Suite (Ascend NPU)

Introduction

The CCE AI Suite (Ascend NPU) add-on supports and manages Huawei NPUs in containers. It provides functions such as automatic driver installation, device registration and scheduling, performance monitoring, and virtual resource management. With this add-on, you can implement automatic deployment, refined scheduling, visualized monitoring, and virtualization of NPUs for diverse heterogeneous computing requirements. It is typically used in:

  • AI model training: It supports multi-NPU parallelism and refined resource scheduling to improve the efficiency and stability of large-scale model training.
  • AI inference and real-time services: It supports low-latency, on-demand computing to ensure real-time performance and service availability of inference tasks.
  • Resource isolation environments: It supports vNPU management to divide resources by granularity for compute isolation and quota control in multi-user scenarios.
  • Training task monitoring and resource optimization: It supports metric collection and visual analysis to monitor training task performance and optimize resource usage.

After this add-on is installed, you can create Ascend-accelerated nodes to quickly and efficiently process inference and image recognition.

A5 HyperNodes support parameter plane RDMA device management. The enable_rdma_shared_dp field in the add-on configuration controls whether to deploy the RDMA Shared Device Plugin. This plugin enables the discovery, allocation, and mounting of RDMA devices on the UB bus, as well as the detection and reporting of NIC faults. For details, see Configuring the RDMA Capability.

Constraints

  • To use Ascend-accelerated nodes in a cluster, the CCE AI Suite (Ascend NPU) add-on must be installed.
  • After an AI-accelerated node is migrated, the node will be reset. If the automatic driver installation function is enabled (supported only by the add-on of version 1.2.5 or later) and the driver corresponding to the NPU node model is selected for the CCE AI Suite (Ascend NPU) add-on of the destination cluster, the NPU driver will be automatically installed after the node migration.
    • If the add-on version is earlier than 1.2.8, you need to manually restart the node after the driver installation for the driver to work.
    • If the add-on version is 1.2.8 or later, the node will be automatically restarted after the driver is installed. After the restart, the driver works.
  • If the NPU driver installation command is included in the post-installation script of the node pool, with automatic driver installation enabled and the correct NPU driver model selected, driver installation commands from the frontend and the npu-driver-installer pod will be executed concurrently to install drivers during node pool scale-out. This may result in discrepancies in the installed driver or installation failures. Therefore, when the driver selection function is enabled for huawei-npu, you are not advised to scale out a node pool that has a compiled post-installation command or post-installation command when creating a node pool to install the NPU driver.
  • Access to Ascend NPUs is controlled by the HwHiAiUser group created during driver installation. The NPU runtime user in the container (whether root or non-root) must belong to this group. Otherwise, npu-smi commands will fail and the NPU cannot be used for inference or training.
  • RDMA device management:
    • RDMA device management is available only in CCE standard and Turbo clusters.
    • The RDMA driver must be installed on any node that uses RDMA functionality, and kernel modules such as ib_uverbs and ib_core must be loaded.
    • The /dev/infiniband/ directory must exist on the node and contain RDMA device files.
    • The RDMA Shared Device Plugin is deployed as a DaemonSet and requires privileged permissions and hostNetwork.
    • RDMA device management supports only Huawei UB bus devices (vendor=0xcc08, deviceID=0x8200).
    • To enable RDMA, set enable_rdma_shared_dp to true. By default, RDMA is disabled.

Installing the Add-on

  1. Log in to the CCE console and click the cluster name to access the cluster console.
  2. In the navigation pane, choose Add-ons. In the right pane, find the CCE AI Suite (Ascend NPU) add-on and click Install.
  3. In the Metric-based Observation area, enable Use NPU-Exporter to Observe NPU Metrics. After this function is enabled, NPU-Exporter will be deployed on the NPU nodes as a DaemonSet. When using this component, pay attention to the following points:

    • NPU-Exporter is supported when the add-on version is 2.1.55 or later. In addition, it must be used with the NPU driver of version 24.x or later.
    • After NPU-Exporter is enabled, if you need to report the collected NPU monitoring data to AOM, see Comprehensive Monitoring of NPU Metrics.

  4. Determine whether to enable Auto Driver Installation (supported only when the add-on version is 1.2.5 or later).

    • Enabled: You can specify the driver version based on the NPU model for easier driver maintenance.
      You can determine whether to enable this function based on different models. After the driver is enabled, the add-on automatically installs the driver based on the specified driver version. By default, the recommended driver is used. You can also select Path to a custom driver from the drop-down list and enter a driver address.
      • The add-on installs the driver based on the driver version selected for the specified model. Such installation is only for nodes with no NPU driver installed. Nodes with an NPU driver installed remain unchanged. If you change the driver version when upgrading the add-on or updating add-on parameters, such change takes effect only on the nodes with no NPU driver installed.
      • After the driver is successfully installed, the node automatically restarts. To prevent service loss during the restart, you are advised to drain the node in advance and then install or upgrade the driver. For details about how to drain a node, see Draining a Node. For details about how to verify the driver installation, see How to Check Whether the NPU Driver Has Been Installed on a Node.
      • Uninstalling the add-on does not automatically delete the installed NPU driver. For details about how to uninstall the NPU driver, see Uninstalling the NPU Driver.
    • Disabled: Driver versions are decided by the system, and the drivers cannot be maintained using the add-on. When you add an NPU node on the console, the system adds the command to install an NPU driver (version and type decided by the system) and automatically restarts the node after the driver installation is complete. Adding an NPU node in another way, such as using an API, requires you to add the driver installation command to Post-installation Script.
    • The following table lists what NPUs and OS specifications are supported.
      Table 1 Specification adaptation

      NPU Type

      Supported OS

      Snt3 (ascend-snt3)

      EulerOS 2.5 x86, CentOS 7.6 x86, EulerOS 2.9 x86, and EulerOS 2.8 Arm

      NOTE:

      The Snt3 Arm model supports up to EulerOS 2.8 Arm, which has now reached EOS. For details, see EOS Plan.

      CCE standard and Turbo clusters v1.28 and later do not support Arm EulerOS 2.8. To use NPUs in these clusters, select compatible NPUs and OSs by referring to Mappings Between Cluster Versions and OS Versions and Software Versions Required by Different Models. For details about the purchase process, see Using Lite Cluster.

  5. Click Install.

Components

Table 2 Add-on components

Component

Description

Resource Type

npu-driver-installer

Used for installing an NPU driver on NPU nodes.

DaemonSet

npu-exporter

Used for monitoring and collecting NPU metric data. You need to manually enable it. After this component is enabled, it runs as a DaemonSet on each NPU node.

DaemonSet

ascend-vnpu-manager

Enables node pool–level NPU virtualization and allows CCE to create vNPUs. By doing so, it facilitates more efficient utilization of NPU resources. After installing the add-on, go to Settings, click the Heterogeneous Resources tab, and enable NPU virtualization. The add-on will then deploy the component on the required nodes. For details, see Static NPU Virtualization.

DaemonSet

NPU Metrics

In CCE standard and Turbo clusters of v1.25 or later, the NPU metrics listed in the table below are exposed through device-plugin. They can be collected, reported, and displayed on AOM. For details about how to collect and monitor more NPU metrics, see Comprehensive Monitoring of NPU Metrics.

Metric

Monitoring Level

Remarks

cce_npu_memory_total

NPU cards

Total NPU memory

cce_npu_memory_used

NPU cards

NPU memory usage

cce_npu_utilization

NPU cards

NPU compute usage

How to Check Whether the NPU Driver Has Been Installed on a Node

After ensuring that the driver is successfully installed, restart the node for the driver to take effect. Otherwise, the driver cannot take effect and NPU resources are unavailable. To check whether the driver is installed, perform the following operations:

  1. On the Add-ons page, click CCE AI Suite (Ascend NPU).

  2. Verify that the node where npu-driver-installer is deployed is in the Running state.

    If the node is restarted before the NPU driver is installed, the driver installation may fail, and a message is displayed on the Nodes page indicating that the driver is not ready. In this case, uninstall the NPU driver from the node and restart the npu-driver-installer pod to reinstall the NPU driver. After confirming that the driver is installed, restart the node. For details about how to uninstall the driver, see Uninstalling the NPU Driver.

Uninstalling the NPU Driver

Log in to the node, obtain the driver operation records in the /var/log/ascend_seclog/operation.log file, and find the driver run package used in the last installation. If the log file does not exist, the driver is installed using the npu_x86_latest.run or npu_arm_latest.run driver combined package. After finding the driver installation package, run the bash {run package name} --uninstall command to uninstall the driver and restart the node as prompted.

  1. Log in to the node where the NPU driver needs to be uninstalled and find the /var/log/ascend_seclog/operation.log file.
  2. If the /var/log/ascend_seclog/operation.log file can be found, view the driver installation log to find the driver installation record.

    If the /var/log/ascend_seclog/operation.log file cannot be found, the driver may be installed using the npu_x86_latest.run or npu_arm_latest.run driver combined package. You can confirm this by checking whether the /usr/local/HiAI/driver/ directory exists.

    The combined package of the NPU driver is stored in the /root/d310_driver directory; other driver installation packages are stored in the /root/npu-drivers directory.

  3. After finding the driver installation package, run the bash {run package path} --uninstall command to uninstall the driver. The following uses Ascend310-hdk-npu-driver_6.0.rc1_linux-x86-64.run as an example:

    bash /root/npu-drivers/Ascend310-hdk-npu-driver_6.0.rc1_linux-x86-64.run --uninstall

  4. Restart the node as prompted. (The installation and uninstallation of the current NPU driver take effect only after the node is restarted.)

Configuring the RDMA Capability

Enabling RDMA

The RDMA capability is controlled by a built-in switch in the add-on.

Parameter

Description

Default Value

enable_rdma_shared_dp

Controls whether to deploy the RDMA Shared Device Plugin.

false

If enable_rdma_shared_dp is set to true, you must also configure the rdma_shared_dp configuration block, which involves the following parameters.

Parameter

Description

Default Value

periodicUpdateInterval

Interval for updating the device status, in seconds.

300

faultDetectPeriod

Fault detection period, in minutes.

5

affinity

Node affinity settings

[{"key":"accelerator/huawei-npu","operator":"Exists"}]

configList[].resourcePrefix

Kubernetes resource name prefix.

huawei.com

configList[].resourceName

Kubernetes resource name.

ub_rdma

configList[].rdmaHcaMax

Maximum number of resource instances that can be allocated to a container from a single physical HCA device.

8

configList[].selectors.buses

Bus type. The value is fixed at ub.

ub

configList[].selectors.vendors

Device vendor ID.

0xcc08

configList[].selectors.deviceIDs

Device ID.

0x8200

The following shows an example configuration:

{
  "periodicUpdateInterval": 300,
  "faultDetectPeriod": 5,
  "configList": [
    {
      "resourcePrefix": "huawei.com",
      "resourceName": "ub_rdma",
      "rdmaHcaMax": 8,
      "selectors": {
        "buses": ["ub"],
        "vendors": ["0xcc08"],
        "deviceIDs": ["0x8200"]
      }
    }
  ]
}

Components

Component

Description

Resource Type

k8s-rdma-shared-dev-plugin

RDMA shared device management add-on, used to discover, allocate, and mount RDMA devices on the UB bus, and detect faults of the 1825 NIC. The health check port is 11257.

DaemonSet

Component Functions

The k8s-rdma-shared-dev-plugin component provides the following functions:

  • Device discovery: automatically discovers RDMA devices on nodes based on configured selectors such as vendor ID, device ID, and bus type.
  • Resource registration: registers RDMA device resources with Kubernetes and supports device sharing mode.
  • Device allocation: responds to device allocation requests from kubelet and allocates RDMA devices to containers.
  • Periodic update: rescans devices and updates the resource list at the interval specified by periodicUpdateInterval.
  • Fault detection: performs fault detection at the interval specified by faultDetectPeriod.
  • Fault reporting: reports device fault information to ConfigMap (dpuinfo-{nodeName}) for the Volcano scheduler to use.

Health Check

Parameter

Value

Check Method

HTTP GET

Check Port

11257

Check Path

N/A

Initial Delay

30 seconds

Check Interval

15 seconds

Timeout

5 seconds

Failure Threshold

3

If the check fails three consecutive times (at most 65 seconds), kubelet automatically restarts the container.

Volume Mounting

Container Mount Path

Host Path

Description

/var/lib/kubelet/device-plugins

/var/lib/kubelet/device-plugins

Communication directory of the kubelet device add-on

/k8s-rdma-shared-dev-plugin

ConfigMap rdma-devices-config

RDMA device configuration file

/dev/

/dev/

Device file directory

/sys

/sys/

sysfs virtual file system, which is used for device enumeration

/dev/infiniband

/dev/infiniband/

InfiniBand device files

/var/log/cce/kubernetes

/var/log/cce/kubernetes

CCE component log directory

/usr/sbin/hinicadm5

/usr/sbin/hinicadm5

1825 NIC management tool (read-only)

/var/log/hinic5

/var/log/hinic5

1825 NIC log directory

Viewing Logs

The RDMA add-on logs are stored in the /var/log/cce/kubernetes/ directory on the node. The log file name is k8s-rdma-shared-dp.log.

Run the following commands to view logs:

# View the pod logs of the RDMA component.
kubectl logs -n kube-system -l k8s-app=rdma-shared-dp-ds

# View the log file on the node.
cat /var/log/cce/kubernetes/k8s-rdma-shared-dp.log

Using a Huawei RDMA Device in a Pod

apiVersion: v1
kind: Pod
metadata:
  name: rdma-test-pod
spec:
  hostNetwork: true  # The RDMA Shared Device Plugin must be used with the host network.
  containers:
  - name: rdma-container
    image: your-rdma-application:latest
    resources:
      requests:
        huawei.com/ub_rdma: '1'
      limits:
        huawei.com/ub_rdma: '1'
  • hostNetwork: true must be configured. Pods can access RDMA devices only through the host network namespace.
  • In the Allocate phase, the add-on automatically mounts necessary RDMA device files (such as /dev/infiniband) into the container. You do not need to manually configure volumeMounts.

Release History

Table 3 CCE AI Suite (Ascend NPU) updates

Add-on Version

Supported Cluster Version

New Feature

2.8.11

v1.29

v1.30

v1.31

v1.32

v1.33

v1.34

v1.35

v1.36

Added support for CANN 8.5.0 in FlexNPU.

2.7.0

v1.29

v1.30

v1.31

v1.32

v1.33

v1.34

v1.35

v1.36

Supported CCE clusters v1.36.

2.5.1

v1.29

v1.30

v1.31

v1.32

v1.33

v1.34

v1.35

Supported CCE clusters v1.35.

2.4.2

v1.28

v1.29

v1.30

v1.31

v1.32

v1.33

v1.34

Supported Snt9b23 devices that do not have RDMA over Converged Ethernet (RoCE) network interfaces.

2.4.1

v1.28

v1.29

v1.30

v1.31

v1.32

v1.33

v1.34

Supported CCE clusters v1.34.

2.3.2

v1.27

v1.28

v1.29

v1.30

v1.31

v1.32

v1.33

Added image signature.

2.2.2

v1.27

v1.28

v1.29

v1.30

v1.31

v1.32

v1.33

Supported CCE clusters v1.33.

2.1.63

v1.25

v1.27

v1.28

v1.29

v1.30

v1.31

v1.32

Supported CCE clusters v1.32.

2.1.53

v1.25

v1.27

v1.28

v1.29

v1.30

v1.31

Fixed the security vulnerabilities.

2.1.46

v1.21

v1.23

v1.25

v1.27

v1.28

v1.29

v1.30

v1.31

Supported CCE clusters v1.31.

2.1.23

v1.21

v1.23

v1.25

v1.27

v1.28

v1.29

v1.30

Fixed some issues.

2.1.22

v1.21

v1.23

v1.25

v1.27

v1.28

v1.29

v1.30

  • Fixed display issues on some pages.
  • Hypernodes can be obtained.
  • NPU topology can be reported.
  • Resolved log printing issues.

2.1.14

v1.21

v1.23

v1.25

v1.27

v1.28

v1.29

v1.30

Fixed some issues.

2.1.7

v1.21

v1.23

v1.25

v1.27

v1.28

v1.29

Resolved the issue that npu-smi fails to be automatically mounted to a service container.

2.1.5

v1.21

v1.23

v1.25

v1.27

v1.28

v1.29

  • Supported CCE clusters v1.29.
  • Added silent fault codes.

2.0.9

v1.21

v1.23

v1.25

v1.27

v1.28

Resolved the issue that process-level fault recovery and adding annotations to workloads occasionally failed.

2.0.5

v1.21

v1.23

v1.25

v1.27

v1.28

  • Supported CCE clusters v1.28.
  • Supported liveness probes.
  • Ascend drivers can be automatically mounted to service containers.

1.2.14

v1.19

v1.21

v1.23

v1.25

v1.27

Supported NPU monitoring.

1.2.6

v1.19

v1.21

v1.23

v1.25

Supported automatic installation of NPU drivers.

1.2.5

v1.19

v1.21

v1.23

v1.25

Supported automatic installation of NPU drivers.

1.2.4

v1.19

v1.21

v1.23

v1.25

Supported CCE clusters v1.25.

1.2.2

v1.19

v1.21

v1.23

Supported CCE clusters v1.23.

1.2.1

v1.19

v1.21

v1.23

Supported CCE clusters v1.23.

1.1.8

v1.15

v1.17

v1.19

v1.21

Supported CCE clusters v1.21.

1.1.2

v1.15

v1.17

v1.19

Added the default seccomp profile.

1.1.1

v1.15

v1.17

v1.19

Supported CCE clusters v1.15.

1.1.0

v1.17

v1.19

Supported CCE clusters v1.19.

1.0.8

v1.13

v1.15

v1.17

Adapted to the Snt3 C75 drivers.

1.0.6

v1.13

v1.15

v1.17

Supported the C75 drivers.

1.0.5

v1.13

v1.15

v1.17

Allowed containers to use Huawei NPUs.

1.0.3

v1.13

v1.15

v1.17

Allowed containers to use Huawei NPUs.