Deploying an AI Application Template
Introduction
AI application templates streamline model deployment via the AI Inference Framework add-on offered by Huawei Cloud CCE.
Traditional deployment of large AI models typically involves complex container configurations, hardware resource (GPU/NPU) scheduling, network gateway settings, and storage mounting. To significantly simplify large model deployment, a one-click inference preset template is now available.
- Preset configurations: The underlying layer includes predefined configurations for mainstream AI models (such as DeepSeek-R1), covering runtime parameters, NPU resource requirements, and hardware adaptation.
- Simplified operations: You can quickly convert a large model into a highly available, low-latency API inference service by simply selecting a template and performing basic configurations through the visualized frontend UI.
Advantages
- Out-of-the-box: No need to edit complex Kubernetes YAML files. Parameters are deeply optimized for Ascend hardware.
- One-click deployment: Complete full lifecycle management, from model loading and resource scheduling to service startup, by simply selecting options on the UI.
- High concurrency and low latency: The underlying layer automatically integrates with high-performance inference engines such as vLLM to provide production-grade inference capabilities.
Prerequisites
Before deployment, ensure the following requirements are met:
- Cluster: A CCE standard or Turbo cluster of v1.28 or later is available.
- Dependent add-on: Volcano Scheduler v1.21.7 or later is installed in the cluster.
- Network access: Inference nodes must have Internet access to pull images and models. For details, see Configuring Internet Access.
Procedure
- Go to the AI Application Templates tab.
- Log in to the CCE console.
- In the navigation pane, choose AI Containers. Then, click the AI Application Templates tab.
- On the AI Application Templates tab page, select the template to be deployed (for example, DeepSeek-R1-Distill-Qwen-1.5B used in this document) and click Deploy.
- On the Deploy Application page, configure parameters.
Table 1 Basic settings Parameter
Description
Application Name
Unique ID of the AI application to be deployed.
AI Application Template
A predefined application template.
The selected template determines the model type (for example, DeepSeek-R1) and deployment architecture (for example, vLLM PD disaggregation).
CAUTION:Switching the application template will change or clear associated data settings. Proceed with caution.
Image
When using a third-party image, ensure the pod can access the Internet. If the third-party image repository is accessible over the public network, pods can pull images directly from the public network. Ensure pods can access the public network using one of the following methods:
NOTE:Ensure sufficient bandwidth when pulling images over the Internet. Insufficient bandwidth may cause slow or failed image pulls.
- In a standard or Turbo cluster, configure an SNAT rule to enable all pods in the cluster to access the Internet. After the configuration, pods can pull images directly over the Internet. For details, see Accessing the Internet from a Container.
- In a standard cluster, bind an EIP to the node running the workload. This allows all workloads on that node to access the Internet. For details, see Binding an EIP.
The vLLM inference engine image (version 0.13.0) optimized for Huawei Ascend NPUs is pre-installed in the environment.
Cluster Name
Select the target cluster for service deployment. The cluster must have the Volcano Scheduler and NPU add-ons pre-installed. Otherwise, deployment will fail.
Namespace
Select the namespace where the inference service will be deployed.
Pods
Specify the number of inference pods to deploy.
- (Optional) Configure a Role. Inference workloads use Roles to organize containers for distributed inference. Each Role contains one entry node and multiple worker nodes.
Table 2 Role settings Parameter
Description
Role Name
Custom identifier, which cannot be modified after the Role is created.
Replicas
Number of instances to create for the Role during deployment.
Entry Container
The entry container is the entry node of a Role. Each Role contains exactly one entry instance, which receives inference requests and coordinates computing tasks for the Role.
- Container Name: defaults to container-entry, identifying the container's responsibilities.
- Resource Quota: If the default configuration does not meet your requirements, click
to adjust CPU, memory, and NPU resources. - Boot Command: defines the main process of the container and executed when the container starts. The main process status determines the container lifecycle. For details, see Startup Command.
- Port: the port used by the container to provide services.
Worker Container
The worker container is the worker node of a Role. Each Role can contain multiple worker instances, which respond to entry scheduling and execute inference tasks.
- Replicas: number of instances for this Role in the same inference service group. Adjust replicas per Role to optimize compute allocation.
- Container Name: defaults to container-worker.
- Container Configuration: Select Same as entry container to reuse all configurations except the container name from the entry container. Select Custom to configure resources, startup commands, and ports independently.
- Boot Command: defines the main process of the container and executed when the container starts. The main process status determines the container lifecycle. For details, see Startup Command.
- Port: the port used by the container to provide services.
- (Optional) Configure advanced settings.
Table 3 Advanced settings Parameter
Description
Custom Inference Metrics
When enabled, Cloud Native Cluster Monitoring automatically collects metrics from inference pods. Required annotations are automatically added to the pod to facilitate metric collection. For details, see Monitoring Custom Metrics Using Cloud Native Cluster Monitoring.
Upgrade Policy
Upgrade mode for the workload:
- Not configured: The upgrade policy is not used.
- Configured:
- Partition: used for grayscale release to control which ServingGroup is upgraded.
If set to 0, all ServingGroups are upgraded. If set to N (N > 0), only ServingGroups with ordinals greater than or equal to N are upgraded.
- Max unavailable pods: maximum number of pods that can be unavailable during the upgrade. It is used to control service availability.
- Partition: used for grayscale release to control which ServingGroup is upgraded.
Scheduling Policies
Supported policies:
- Gang scheduling: ensures that a group of pods (for example, multiple Roles of a distributed job) are scheduled together or not at all.
- Not configured: Gang scheduling is not used.
- Min Role replicas: defines the minimum number of Role replicas. The scheduler schedules pods only when sufficient resources are available.
- Topology-aware Scheduling
- Not configured: Topology-aware scheduling is not used.
- Configured: Select inference pod affinity or Role affinity.
- Pod Affinity Tier: affinity tier for ServingGroup pod scheduling. For example, deploy a PD disaggregation group within the same network domain.
- Role Affinity Tier: affinity tier for Role pod scheduling. For example, deploy distributed inference pods (entry and workers) within the same network domain.
Tolerance Policies
Using both taints and tolerations allows (but does not force) the pod to be scheduled to a node with matching taints, and controls pod eviction policies after the node is tainted. For details, see Configuring Tolerance Policies.
Labels and Annotations
Add labels or annotations to workload pods as key-value pairs. For details, see Configuring Labels and Annotations.
Description
Description of the inference workload.
- Confirm the configurations and click Submit.
On the Inference Workloads page, view the new inference workload in the Deploying state.
- Configure a Service to expose the deployed model. This example uses NodePort access. In the selector, set modelserving.volcano.sh/name to the inference workload name and modelserving.volcano.sh/role to proxy. Set both the container port and the service port to 8181. For details, see NodePort.

- Verify the model by sending a standard POST request to the port exposed by the Service.
curl -X POST http://<node-IP>:<node-port>/v1/chat/completions \ -H "Content-Type: application/json" \ -d '{ "model": "ds_r1", "messages": [ { "role": "user", "content": "Hello, how are you?" } ], "max_tokens": 100 }'If the model is running correctly, the response returns JSON data containing choices and message fields, where the content field holds the model's generated reply.
{ "id": "chatcmpl-753ceb4c-7aa5-4a4f-94b5-dc791c*****", "object": "chat.completion", "created": 1779936213, "model": "ds_r1", "choices": [ { "index": 0, "message": { "role": "assistant", "content": "Alright, someone just said \"Hello, how are you?\" I should respond in a friendly and approachable way.\n\nI need to keep it simple and open-ended to encourage them to share more.\n\nMaybe ask them how they're doing or if they have any questions they want to discuss.\n\nThat should make the conversation feel natural and helpful.\n</think>\n\nHello! I'm just a computer program, so I don't have feelings, but thanks for asking! How can I assist you today?", "refusal": null, "annotations": null, "audio": null, "function_call": null, "tool_calls": [], "reasoning": null, "reasoning_content": null }, "logprobs": null, "finish_reason": "stop", "stop_reason": null, "token_ids": null } ], "service_tier": null, "system_fingerprint": null, "usage": { "prompt_tokens": 11, "total_tokens": 109, "completion_tokens": 98, "prompt_tokens_details": null }, "prompt_logprobs": null, "prompt_token_ids": null, "kv_transfer_params": null }
What is your overall rating for this page?
Thank you very much for your feedback. We will continue working to improve the documentation.See the reply and handling status in My Cloud VOC.
For any further questions, feel free to contact us through the chatbot.
Chatbot