ALM-18009 Heap Memory Usage of JobHistoryServer Exceeds the Threshold
Alarm Description
The system checks the heap memory usage of MapReduce JobHistoryServer every 30 seconds and compares the actual usage with the threshold. This alarm is generated when the heap memory usage of MapReduce JobHistoryServer exceeds the threshold.
Users can choose O&M > Alarm > Thresholds > Name of the desired cluster > MapReduce to change the threshold.
When the Trigger Count is 1, this alarm is cleared when the heap memory usage of MapReduce JobHistoryServer is less than or equal to the threshold. When the Trigger Count is greater than 1, this alarm is cleared when the heap memory usage of MapReduce JobHistoryServer is less than or equal to 95% of the threshold.
Alarm Attributes
| Alarm ID | Alarm Severity | Auto Cleared |
|---|---|---|
| 18009 | Critical (default threshold: 95% of the maximum memory) Major (default threshold: 90% of the maximum memory) | Yes |
Alarm Parameters
| Parameter | Description |
|---|---|
| Source | Specifies the cluster for which the alarm was generated. |
| ServiceName | Specifies the service for which the alarm was generated. |
| RoleName | Specifies the role for which the alarm was generated. |
| HostName | Specifies the host for which the alarm was generated. |
| Trigger Condition | Specifies the threshold for triggering the alarm. |
Impact on the System
When the heap memory usage of MapReduce JobHistoryServer is overhigh, the performance of MapReduce log archiving is affected. What is more, a memory overflow occurs so that the YARN service is unavailable.
Possible Causes
The heap memory of the MapReduce JobHistoryServer instance on the node is overused or the heap memory is inappropriately allocated. As a result, the usage exceeds the threshold.
Handling Procedure
Check the memory usage.
- On Manager, choose O&M > Alarm > Alarms to view details about the current alarm and record the host name of the instance for which the alarm is generated.
For details about how to log in to FusionInsight Manager, see Accessing MRS Manager.
- On Manager, choose Cluster > Services > Mapreduce and click the Instances tab. On the displayed page, select JobHistoryServer (the host name of the instance for which the alarm is generated) and select Customize > Resource from the drop-down list in the upper right corner of the Chart area. On the displayed page, select JobHistoryServer Heap Memory Usage Statistics. Check the heap memory usage.
- Check whether the used heap memory of JobHistoryServer reaches 95% of the maximum heap memory specified for JobHistoryServer.
- On Manager, choose Cluster > Services > Mapreduce > Configurations > All Configurations > JobHistoryServer > System. Increase the value of GC_OPTS parameter as required, click Save. Click OK to restart the JobHistoryServer.
- The mapping between the number of historical tasks (10000) and the memory of the JobHistoryServer is as follows:
- During the restart of JobHistoryServer, the status of tasks such as Hive cannot be queried. As a result, the query result may be inaccurate.
- Check whether the alarm is cleared 5 minutes later.
- If yes, no further action is required.
- If no, go to Step 6.
Collect fault information.
- On Manager, choose O&M > Log > Download.
- In the Service area, select the following nodes of the desired cluster.
- NodeAgent
- MapReduce
- Click
in the upper right corner, and set Start Date and End Date for log collection to 10 minutes before and after the alarm generation time, respectively. Then, click Download. - Send the collected fault logs to O&M personnel for help.
Alarm Clearance
This alarm is automatically cleared after the fault is rectified.
Related Information
None
What is your overall rating for this page?
Thank you very much for your feedback. We will continue working to improve the documentation.See the reply and handling status in My Cloud VOC.
For any further questions, feel free to contact us through the chatbot.
Chatbot