Updated on 2026-08-27 GMT+08:00

Overview

Distributed Training

Distributed training speeds up deep learning by running tasks simultaneously across several compute nodes like servers or GPUs, enabling faster model training or handling bigger datasets. Distributed training splits a training task across multiple nodes, where each node processes a portion of the model. These nodes share their computed results through a communication system, enabling the full model to be trained efficiently. This approach greatly enhances training speed, particularly for complex models and large datasets.

ModelArts enables distributed training by automatically managing node communications and resources for effective parallel processing. This section describes the functions of DDP for multi-node multi-PU distributed training.

Constraints

  • If the notebook instance flavors are changed, you can only perform single-node debugging. You cannot perform distributed debugging or submit remote training jobs.
  • When using Ascend accelerators for preset specification training in a distributed, non-8-PU scenario, you are advised to use the auto-negotiation method for link establishment.
  • Only the PyTorch and MindSpore AI frameworks can be used for multi-node distributed debugging. If you want to use MindSpore, each node must be equipped with eight PUs.
  • The OBS paths in the debugging code should be replaced with your OBS paths.
  • PyTorch is used to write debugging code in this document. The process is the same for different AI frameworks. You only need to modify some parameters.

Billing

Model training in ModelArts uses compute and storage resources, which are billed. Compute resources are billed for running training jobs. Storage resources are billed for storing data in OBS or SFS. For details, see Model Training Billing Items.

Advantages of Multi-Node Multi-PU Training Using DistributedDataParallel

  • Fast communication
  • Balanced load
  • Fast running speed

Getting Started