Help Center/ ModelArts/ Troubleshooting/ Training Jobs/ Training Performance Issues/ Training Performance Deteriorated
Updated on 2024-06-11 GMT+08:00
Training Performance Deteriorated
Symptom
When a ModelArts algorithm is used for training, it will take more time than expected for training.
Possible Causes
The possible causes are as follows:
- The job code or training parameters have been modified.
- The GPU hardware for training malfunctions.
Solution
- Check whether the training code and parameters have been modified.
- Check whether the allocation of the CPU, memory, GPU, snt9, or Infiniband resources complies with the expectation.
- Use CloudShell to log in to the Linux and check the GPU working status.
- Run the nvidia-smi command to check whether the GPU is working properly.
- Run the nvidia-smi -q -d TEMPERATURE command to check the temperature. If the temperature is too high, the training performance deteriorates.
Parent topic: Training Performance Issues
What is your overall rating for this page?
0
1
2
3
4
5
6
7
8
9
10
Very dissatisfiedVery satisfied
Thank you very much for your feedback. We will continue working to improve the documentation.
The system is busy. Please try again later.