ALM-14040 Number of Slow SyncWriterOsCache on HDFS DataNodes Per Second Exceeds the Threshold
Alarm Description
The system periodically checks the number of slow SyncWriterOsCache occurrences per second on HDFS DataNode instances every 60 seconds and compares it with the threshold. This alarm is generated when the number of slow SyncWriterOsCache occurrences per second exceeds the threshold for 3 minutes.
This alarm is cleared when the number of slow SyncWriterOsCache occurrences per second is less than or equal to the threshold.
This section applies to MRS 3.6.0-LTS and later versions.
Alarm Attributes
| Alarm ID | Alarm Severity | Auto Cleared |
|---|---|---|
| 14040 | Major (default threshold: 100) | Yes |
Alarm Parameters
| Type | Parameter | Description |
|---|---|---|
| Location Information | Source | Specifies the cluster for which the alarm was generated. |
| ServiceName | Specifies the service for which the alarm was generated. | |
| RoleName | Specifies the role for which the alarm was generated. | |
| HostName | Specifies the host for which the alarm was generated. | |
| Additional Information | Trigger Condition | Specifies the alarm triggering condition. |
Impact on the System
Slow SyncWriterOsCache in HDFS impacts its data read and write performance.
Possible Causes
- The alarm threshold is improperly set.
- The HDFS DataNode's processing capacity has reached a bottleneck.
Handling Procedure
Check whether the alarm threshold is set properly.
- Log in to FusionInsight Manager and choose O&M > Alarm > Alarms. In the Location field of the alarm details, view the host name of the DataNode instance for which this alarm is generated.
- Choose Cluster > Services > HDFS, click the Instances tab, and click the DataNode role based on the host name obtained in Step 1.
- Choose Chart > Performance, view the Slow SyncWriterOsCache Occurrences Per Second chart, and obtain the peak value within one day before and after the alarm is generated.
- Choose O&M > Alarm > Thresholds, locate the desired cluster, choose HDFS, and click Slow SyncWriterOsCache Occurrences Per Second. Click Modify in the Operation column of the default rule, and change the threshold to 150% of the peak value displayed within one day before and after the alarm is generated. Click OK to save the new threshold.
- Wait 5 minutes and check whether the alarm is cleared.
- If yes, no further action is required.
- If no, go to Step 6.
Check whether the DataNode's processing capability reaches a bottleneck.
- On FusionInsight Manager, choose O&M > Alarm > Alarms to check whether the "ALM-14015 DataNode GC Time Exceeds the Threshold" alarm is reported and whether the host for which the alarm is generated is the same as that in Step 1.
- Rectify the fault by following the handling procedure of "ALM-14015 DataNode GC Time Exceeds the Threshold".
- Wait 5 minutes and check whether the alarm is cleared.
- If yes, no further action is required.
- If no, go to Step 9.
Collect fault information.
- On FusionInsight Manager, choose O&M. In the navigation pane on the left, choose Log > Download.
- Expand the Service drop-down list, and select HDFS for the target cluster.
- Click the edit icon in the upper right corner, and set Start Date and End Date for log collection to 10 minutes ahead of and after the alarm generation time, respectively. Then, click Download.
- Contact O&M engineers and provide the collected logs.
Alarm Clearance
This alarm is automatically cleared after the fault is rectified.
Related Information
None
What is your overall rating for this page?
Thank you very much for your feedback. We will continue working to improve the documentation.See the reply and handling status in My Cloud VOC.
For any further questions, feel free to contact us through the chatbot.
Chatbot