Introduction
Overview
Cross-cloud AI compute collaboration can aggregate compute, data, and ecosystems, provide super compute for foundation model training, promote the implementation of AI algorithm development in the industry, activate industry integration and co-evolution, achieve green development, and improve energy utilization efficiency.
Compute is scheduled across clouds. Compute enablement is preferentially carried out on clouds. Cloud services are used to schedule various AI compute resources (NPUs/GPUs). The unified scheduling of large-scale, cross-domain, and heterogeneous compute is used to manage and schedule AI compute resources across sovereign clouds in a unified manner.
Each sovereign cloud independently builds compute enablement and independently connects to inter-cloud collaboration to schedule compute across sovereign clouds. Sovereign cloud vendors, as compute providers, do not depend on each other and can flexibly access and exit.
Functions
- Operations collaboration: AI compute resources are globally visible. Accounts are streamlined. A multi-tenant user can use AI compute resources of multiple sovereign clouds on the local cloud.
- Scheduling collaboration: A multi-tenant user can centrally schedule different training jobs from the local cloud to different intelligent computing centers on the collaboration cloud. This allows jobs and data to be forwarded from one intelligent computing center to another, enabling intelligent scheduling of jobs and data across AI computing centers. This ensures load balancing of compute across the entire network and maximizes resource utilization.
- Collaborative training across sovereign clouds: A multi-tenant user can schedule sub-tasks of a training job to multiple intelligent computing centers on the collaboration cloud for collaborative training based on the partitioning strategy on the local cloud. Sub-tasks of a single big computing job are distributed and parallel across intelligent computing centers. This solves the problem that a single intelligent computing center cannot meet the compute requirements of foundation model training, enabling free flow of compute.
- Training performance loss of multiple intelligent computing centers: The iteration time performance gap between cross-cluster and single-cluster training jobs is within 10%.
- Expansion efficiency of multiple intelligent computing centers: Compared with independent scheduling of a single intelligent computing center, coordinated scheduling across intelligent computing centers improves the execution efficiency of training computing tasks by 80%, enhancing the overall resource utilization of intelligent computing centers.
Concepts
- Local cloud: A user's main workloads are retained on the local cloud. After the local cloud is interconnected with the collaboration cloud, the user can request collaborative cloud services, AI compute resources, and training job provisioning on the local cloud to perform operations such as cross-cloud collaborative training.
- Collaboration cloud: Based on open collaboration between different sovereign clouds, after the collaboration cloud is interconnected with the local cloud, a user's workloads on the local cloud are diverted to the collaboration cloud, maximizing the compute utilization of the collaboration cloud.
What is your overall rating for this page?
Thank you very much for your feedback. We will continue working to improve the documentation.See the reply and handling status in My Cloud VOC.
For any further questions, feel free to contact us through the chatbot.
Chatbot