ALM-50842 Percentage of Concurrent StoreWorker Writes Exceeds the Threshold
Alarm Description
The system checks the percentage of concurrent writes on StoreWorker instances every 30 seconds. This alarm is generated when the percentage exceeds the threshold.
This alarm is cleared when the percentage of concurrent writes on StoreWorker instances falls below the threshold.
Alarm Attributes
| Alarm ID | Alarm Severity | Auto Cleared |
|---|---|---|
| 50842 | Minor | Yes |
Alarm Parameters
| Parameter | Description |
|---|---|
| Source | Specifies the cluster or system for which the alarm was generated. |
| ServiceName | Specifies the service for which the alarm was generated. |
| RoleName | Specifies the role for which the alarm was generated. |
| HostName | Specifies the host for which the alarm was generated. |
Impact on the System
Too many partitions on a node are executing shuffle writes. As a result, the server efficiency decreases.
Possible Causes
- The load tilt is too large, and the load pressure is too high.
- The number of nodes is too small, causing high pressure on each node.
Handling Procedure
Check the percentage of nodes for which the alarm is generated.
- On FusionInsight Manager, choose O&M > Alarm > Alarms. In the alarm list, view the role name and obtain the IP address of the instance in Location of the alarm whose ID is 50842.
- Choose Cluster > Services > MemArtsStore > Instances, and calculate the percentage of the StoreWorker instances for which the alarm is generated in all StoreWorker instances.
Add nodes.
- Choose Cluster > Services > MemArtsStore and click Instances.
- Click Add Instance to add StoreWorker instances.
- Wait 2 minutes and check whether the alarm is automatically cleared.
- If yes, no further action is required.
- If no, go to Step 8.
Change the Spark parameter value to the native setting.
- Log in to the Spark client node, go to the /opt/huawei/OneWork/client/Spark/spark/conf directory, change the value of spark.shuffle.manager in the spark-defaults.conf file to sort, save the modification, exit, and run the task again.
spark.shuffle.manager = sort
- Wait 2 minutes and check whether the alarm is automatically cleared.
- If yes, no further action is required.
- If no, go to Step 8.
Collect fault information.
- On FusionInsight Manager, choose O&M. In the navigation pane on the left, choose Log > Download.
- Expand the Service drop-down list, and select MemArtsStore for the target cluster.
- Click the edit icon in the upper right corner, and set Start Date and End Date for log collection to 10 minutes ahead of and after the alarm generation time, respectively. Then, click Download.
- Contact O&M engineers and send the collected logs.
Alarm Clearance
This alarm is automatically cleared after the fault is rectified.
Related Information
None
What is your overall rating for this page?
Thank you very much for your feedback. We will continue working to improve the documentation.See the reply and handling status in My Cloud VOC.
For any further questions, feel free to contact us through the chatbot.
Chatbot