ALM-14045 Average ReadBlock Operation Time Exceeds the Threshold
Alarm Description
The system checks the average time of the ReadBlock operation on the DataNode every 30 seconds. This alarm is generated when the average time exceeds the threshold for a specified number of times (20 times by default) consecutively.
This alarm is cleared when the average time of the ReadBlock operation on the DataNode falls below the threshold.
This section applies to MRS 3.6.0-LTS.1 and later versions.
Alarm Attributes
| Alarm ID | Alarm Severity | Auto Cleared |
|---|---|---|
| 14045 | Minor (default threshold: 5,000 ms) Major (default threshold: 10,000 ms) | Yes |
Alarm Parameters
| Type | Parameter | Description |
|---|---|---|
| Location Information | Source | Specifies the cluster for which the alarm was generated. |
| ServiceName | Specifies the service for which the alarm was generated. | |
| RoleName | Specifies the role for which the alarm was generated. | |
| HostName | Specifies the host for which the alarm was generated. | |
| Additional Information | Trigger Condition | Specifies the alarm triggering condition. |
Impact on the System
If block reading on the DataNode becomes slow, services that depend on HDFS for data reading become slow.
Possible Causes
- The alarm threshold is improperly set.
- The memory allocated to the DataNode is insufficient and frame freezing occurs on the JVM due to frequent Full GCs.
- The disk of the node where the DataNode is located is slow in performance.
- The disk write speed of the DataNode OS is slow.
- The network between the client and the DataNode is faulty.
Handling Procedure
Check whether the alarm threshold is proper.
- Log in to MRS Manager, choose O&M > Alarm > Alarms, view the alarm details, and obtain the host name of the DataNode for which the alarm is generated.
- Check whether the services that depend on HDFS are running normally and whether the services are running slowly or timed out.
- On Manager, choose Cluster > Services > HDFS > Instances, click the DataNode role name of the host name obtained in Step 1, choose Chart > Operation, and check the indicator DataNode Operation Time to obtain its peak value within one day before and after the alarm is generated.
- Choose O&M > Alarm > Thresholds > Name of the desired cluster > HDFS, locate the Average ReadBlock Operation Time, click Modify in the Operation column of the default rule, change the threshold to 150% of the peak value of the indicator within one day before and after the alarm is generated, and click OK.
- Wait 5 minutes and check whether the alarm is automatically cleared.
- If yes, no further action is required.
- If no, go to Step 6.
Check whether the memory configured for the DataNode is proper.
- On Manager, choose O&M > Alarm > Alarms to check whether the ALM-14015 DataNode GC Time Exceeds the Threshold alarm is generated and whether the host of the alarm is the same as that obtained in Step 1.
- Click View Help in the row where the alarm is located and handle the alarm by referring to the help document.
- After the ALM-14015 alarm is cleared, wait for 10 minutes and check whether the alarm is automatically cleared.
- If yes, no further action is required.
- If no, go to Step 9.
Check whether the disk on the node where the DataNode is located is slow in performance.
- On Manager, choose O&M > Alarm > Alarms and check whether there is any of the following alarms: ALM-12180 Suspended Disk I/O, ALM-12191 Disk I/O Usage Exceeds the Threshold, ALM-12204 Wait Duration of a Disk Read Exceeds the Threshold, and ALM-12205 Wait Duration of a Disk Write Exceeds the Threshold, and whether the host of the alarm is the same as that obtained in Step 1.
- Click View Help in the row where the alarm is located and handle the alarm by referring to the help document.
- Wait for 10 minutes and check whether the alarm is automatically cleared.
- If yes, no further action is required.
- If no, go to Step 12.
Check whether the disk write speed of the DataNode OS is slow.
- On Manager, choose Cluster > Services > HDFS > Instances, click the DataNode role name corresponding to the host name obtained in Step 1, and choose Chart > Performance. Check whether the indicators Number of Slow IOs Per Second, Slow WriteDataToDisk Occurrences Per Second, and Slow Flush or Sync Occurrences Per Second are abnormal during the period when the alarm is generated.
- If yes, contact O&M engineers to check the disk performance.
- If no, go to Step 14.
- Wait 5 minutes and check whether the alarm is automatically cleared.
- If yes, no further action is required.
- If no, go to Step 14.
Check whether the network between DataNodes is slow.
- On Manager, choose Cluster > Services > HDFS > Instances, click the DataNode role name corresponding to the host name obtained in Step 1, and choose Chart > Performance. Check the indicator Slow WritePacketToDownStream Occurrences Per Second. Check whether the indicator is abnormal when the alarm is generated.
- If yes, contact O&M engineers to rectify the network fault.
- If no, go to Step 16.
- Wait 5 minutes and check whether the alarm is automatically cleared.
- If yes, no further action is required.
- If no, go to Step 16.
Collect fault information.
- On MRS Manager, choose O&M. In the navigation pane on the left, choose Log > Download.
- Expand the Service drop-down list, select the HDFS services for the target cluster, and click OK.
- Click the edit icon in the upper right corner, and set Start Date and End Date for log collection to 10 minutes ahead of and after the alarm generation time, respectively. Then, click Download.
- Send the collected fault logs to O&M engineers for help.
Alarm Clearance
This alarm is automatically cleared after the fault is rectified.
Related Information
None
What is your overall rating for this page?
Thank you very much for your feedback. We will continue working to improve the documentation.See the reply and handling status in My Cloud VOC.
For any further questions, feel free to contact us through the chatbot.
Chatbot