AI Container Overview
AI Container is an AI workload scheduling platform built on CCE, designed for large-scale AI development, training, and inference. As large language models (LLMs) continue to grow in parameter scale, AI training and inference jobs place unique demands on infrastructure: heterogeneous compute scheduling, distributed training communication optimization, and high-concurrency inference service governance. Traditional container schedulers lack the specialized capabilities to address these requirements. Built on CCE container clusters and powered by Volcano (a high-performance workload scheduling engine), AI Container addresses these challenges through resource management, monitoring and maintenance, performance acceleration, workload orchestration and scheduling, and intelligent auto scaling, enabling enterprises to rapidly build cloud-native AI infrastructure.
AI Container is ideal for large-scale LLM distributed training, Prefill-Decode (PD) inference deployment, and multi-tenant GPU/NPU cluster management.
Advantages
- Prebuilt AI application templates for rapid deployment
Out-of-the-box, scenario-specific configurations eliminate complex parameter tuning, streamline the path from model deployment to inference, and accelerate AI application development.
- Unified resource management for AI operations
Unified management of AI infrastructure resources, including GPUs, NPUs, Remote Direct Memory Access (RDMA), high-performance caches, and distributed storage, with built-in monitoring and fault detection for AI workloads.
- Intelligent scheduling to reduce costs
Based on the characteristics of AI computing tasks, AI Container provides scheduling policies including gang scheduling, bin packing, topology-aware scheduling, GPU virtualization scheduling, and NPU compact scale-in. These policies reduce resource fragmentation and improve cluster resource utilization.
Capability Overview
The table below lists the core capabilities of AI Container.
| Capability | Description |
|---|---|
| Deployment and management of AI model inference services, including inference frameworks, inference gateways, and advanced architectures such as PD disaggregation. | |
| Distributed training job management via Volcano jobs, with queue-based resource allocation and priority scheduling across multiple tenants. | |
| Out-of-the-box, scenario-specific templates with one-click deployment, streamlining the path from model development to inference. |
What is your overall rating for this page?
Thank you very much for your feedback. We will continue working to improve the documentation.See the reply and handling status in My Cloud VOC.
For any further questions, feel free to contact us through the chatbot.
Chatbot