HA Configuration Scenarios
This section recommends HA configurations for common scenarios, providing a reference for setting up your HA training jobs.
| Scenario | Description | Recommended Configuration | Configuration Combination | Recommendation |
|---|---|---|---|---|
| Single-node training | Suitable for single-node single-PU and single-node multi-PU training. |
| Resumable training | Recommended |
| Auto restart | Recommended | |||
| Job suspension detection | Optional | |||
| Restart upon suspension | Optional | |||
| Operator re-execution | Usually not needed | |||
| Multi-node multi-PU training | Suitable for distributed training and large-scale GPU/NPU training. |
| Resumable training | Required |
| Auto restart | Recommended | |||
| Unconditional auto restart | Recommended | |||
| Job suspension detection | Recommended | |||
| Restart upon suspension | Recommended | |||
| Pod-level rescheduling | Based on resource pool capabilities | |||
| Isolated job-level rescheduling | Based on resource pool capabilities | |||
| LLM long-term stable training | Suitable for tasks with long durations and large resource scales, such as LLM pre-training, fine-tuning, and reinforcement learning. |
| Resumable training | Required |
| Auto restart | Required | |||
| Unconditional auto restart | Recommended | |||
| Suspension detection | Recommended | |||
| Restart upon suspension | Recommended | |||
| Operator re-execution | Recommended for specific Ascend supernode scenarios | |||
| Log failure analysis | Recommended |
What is your overall rating for this page?
Thank you very much for your feedback. We will continue working to improve the documentation.See the reply and handling status in My Cloud VOC.
For any further questions, feel free to contact us through the chatbot.
Chatbot