Updated on 2026-08-13 GMT+08:00

AI Container Overview

AI Container is an AI workload scheduling platform built on CCE, designed for large-scale AI development, training, and inference. As large language models (LLMs) continue to grow in parameter scale, AI training and inference jobs place unique demands on infrastructure: heterogeneous compute scheduling, distributed training communication optimization, and high-concurrency inference service governance. Traditional container schedulers lack the specialized capabilities to address these requirements. Built on CCE container clusters and powered by Volcano (a high-performance workload scheduling engine), AI Container addresses these challenges through resource management, monitoring and maintenance, performance acceleration, workload orchestration and scheduling, and intelligent auto scaling, enabling enterprises to rapidly build cloud-native AI infrastructure.

AI Container is ideal for large-scale LLM distributed training, Prefill-Decode (PD) inference deployment, and multi-tenant GPU/NPU cluster management.

Advantages

  • Prebuilt AI application templates for rapid deployment

    Out-of-the-box, scenario-specific configurations eliminate complex parameter tuning, streamline the path from model deployment to inference, and accelerate AI application development.

  • Unified resource management for AI operations

    Unified management of AI infrastructure resources, including GPUs, NPUs, Remote Direct Memory Access (RDMA), high-performance caches, and distributed storage, with built-in monitoring and fault detection for AI workloads.

Capability Overview

The table below lists the core capabilities of AI Container.

Capability

Description

Inference workload

Deployment and management of AI model inference services, including inference frameworks, inference gateways, and advanced architectures such as PD disaggregation.

Training job

Distributed training job management via Volcano jobs, with queue-based resource allocation and priority scheduling across multiple tenants.

AI application template

Out-of-the-box, scenario-specific templates with one-click deployment, streamlining the path from model development to inference.