ALM-38009 Kafka Topic Overload (Applicable to MRS 3.1.0 and Earlier Versions)
If the alarm is "ALM-38009 Busy Broker Disk I/Os", rectify the fault by referring to ALM-38009 Busy Broker Disk I/Os (Applicable to Versions Later Than MRS 3.1.0).
Alarm Description
The system checks the overload status of each Kafka topic every 60 seconds. This alarm is generated when the percentage of partitions of a topic on the overloaded disk exceeds the threshold (40% by default).
Its Trigger Count is 1. This alarm is cleared when the percentage of partitions of a topic on the overloaded disk is lower than the threshold (40% by default).
An overloaded disk refers to the disk whose I/O usage of a disk partition is greater than 80%.
Example:
The partitions of Topic A are distributed on three brokers. The I/O usages of the disk partitions on two brokers are greater than 80%.
The percentage of partitions on the overloaded disk is 2/3, greater than 40%, thus this alarm is generated.
This section applies only to MRS 3.1.0 or earlier.
Alarm Attributes
| Alarm ID | Alarm Severity | Auto Cleared |
|---|---|---|
| 38009 | Major | Yes |
Alarm Parameters
| Parameter | Description |
|---|---|
| Source | Specifies the cluster for which the alarm was generated. |
| ServiceName | Specifies the service for which the alarm was generated. |
| RoleName | Specifies the role for which the alarm was generated. |
| HostName | Specifies the host for which the alarm was generated. |
| TopicName | Specifies the Kafka topic for which the alarm was generated. |
Impact on the System
The disk partition has frequent I/Os. Data may fail to be written to the Kafka topic for which the alarm is generated.
Possible Causes
- There are many replicas configured for the topic.
- The parameter for batch writing producer's messages is inappropriately configured. The service traffic of this topic is too heavy, and the current partition configuration is inappropriate.
Handling Procedure
Check the number of topic replicas.
- On FusionInsight Manager, choose O&M > Alarm > Alarms. Locate the row that contains this alarm, click
, and view the host name in Location. - On FusionInsight Manager, choose Cluster > Services > Kafka > KafkaTopic Monitor, search for the topic for which the alarm is generated, and check the number of replicas.
- Reduce the replication factors of the topic (for example, reduce to 3) if the number of replicas is greater than 3.
Run the following command on the cluster client to replan the replicas of Kafka topics:
kafka-reassign-partitions.sh --zookeeper {zk_host}:{port}/kafka --reassignment-json-file {manual assignment json file path} --execute
For details about operations on the Kafka client, see Using the Kafka Client.
Example:
/opt/client/Kafka/kafka/bin/kafka-reassign-partitions.sh --zookeeper 10.149.0.90:2181,10.149.0.91:2181,10.149.0.92:2181/kafka --reassignment-json-file expand-cluster-reassignment.json --executeIn the expand-cluster-reassignment.json file, specify the target brokers to which the topic's partitions will be migrated. The JSON content should follow this format:
{"partitions":[{"topic": "topicName","partition": 1,"replicas": [1,2,3] }],"version":1}. - Check whether the alarm is cleared after a period of time.
- If yes, no further action is required.
- If no, go to Step 5.
Check the partition planning of the topic.
- On the KafkaTopic Monitor page, choose Topic Traffic > Topic Input Traffic of each topic to obtain the topic with the largest value of Topic Input Traffic, and check partitions on this topic and information about hosts of these partitions.
- Log in to the host queried in Step 5 and run the following command to check the %util value of each disk.
iostat -d -x
- If the %util value of each disk exceeds the threshold (80% by default), expand the Kafka disk capacity and replan the topic partition by referring to Step 3.
- If the %util values of the disks vary greatly, check the disk partition configuration of Kafka.
For example, check the value of log.dirs in the ${BIGDATA_HOME}/FusionInsight_HD_*/1_14_Broker/etc/server.properties file.
Run the following command to view the Filesystem information:
df -h log.dirs valueThe command output is as follows.

- If the partition where Filesystem is located matches the partition with a high %util value, plan Kafka partitions on idle disks, configure log.dirs as an idle disk directory, and replan topic partitions by referring to Step 3. Ensure that the partitions of the topic are evenly distributed to each disk.
- Check whether the alarm is cleared after a period of time.
- Check whether the alarm is cleared after a period of time.
- If yes, no further action is required.
- If no, go to Step 9.
Collect fault information.
- On FusionInsight Manager, choose O&M. In the navigation pane on the left, choose Log > Download.
- Expand the Service drop-down list, and select Kafka for the target cluster.
- Click
in the upper right corner, and set Start Date and End Date for log collection to 10 minutes before and after the alarm generation time, respectively. Then, click Download. - Send the collected logs to O&M engineers.
Alarm Clearance
This alarm is automatically cleared after the fault is rectified.
Related Information
None
What is your overall rating for this page?
Thank you very much for your feedback. We will continue working to improve the documentation.See the reply and handling status in My Cloud VOC.
For any further questions, feel free to contact us through the chatbot.
Chatbot