
# 使用PrometheusRule配置普罗监控与告警规则
Prometheus具有PrometheusRule的能力，PrometheusRule提供了一种用于监控和警报的规则语言，能够方便用户更好地使用Prometheus查询监控指标，配置基于PromQL的告警规则。
![](https://support.huaweicloud.com/usermanual-cce/public_sys-resources/note_3.0-zh-cn.png)
当前云原生监控插件仅在开启本地数据存储时，支持PrometheusRule配置。
#### 通过PrometheusRule配置记录规则
1. Prometheus提供了PrometheusRule用于创建用户自己的record来查询指标。
   ```
   apiVersion: monitoring.coreos.com/v1
   kind: PrometheusRule
   metadata:
     name: recording-rules-demo
     namespace: monitoring
     labels:
       role: operator-prometheus   # 保持一致，必须配置，prometheus配置了该ruleSelector
   spec: 
     groups: 
     - name:  demo
       interval: 15s
       rules:
       - record: cpu_request
         expr:   kube_pod_container_resource_requests{resource="cpu",unit="core"}
       - record: cpu_limit
         expr:   kube_pod_container_resource_limits{resource="cpu",unit="core"}
       - record: memory_request
         expr:   kube_pod_container_resource_requests{resource="memory",unit="byte"}
       - record: memory_limit
         expr:   kube_pod_container_resource_limits{resource="memory",unit="byte"}
   ```
   
2. 创建成功后，可以访问Prometheus的Web页面，在"Status \> Rules"页面中找到配置的PrometheusRule。 ![](https://support.huaweicloud.com/usermanual-cce/zh-cn_image_0000001756896281.png "点击放大")
   
 
#### 通过PrometheusRule配置告警规则
通过配置PrometheusRule的CR资源来创建普罗的告警规则。本文以集群CPU使用率告警为例创建告警配置模板。
1. 创建示例的告警规则模板。
   ```
   kubectl apply -f PrometheusRule.yaml
   ```
   PrometheusRule.yaml文件内容如下：
   ```
   apiVersion: monitoring.coreos.com/v1
   kind: PrometheusRule
   metadata:
     labels:
       role: operator-prometheus # 保持一致，必须配置，prometheus配置了该ruleSelector
     name: alert-rules-demo
     namespace: monitoring
   spec:
     groups:
     - name: alert-cluster-demo
       rules:
       - alert: 集群CPU使用率超过50%
         expr: 100 - (avg  (irate(node_cpu_seconds_total{mode="idle"}[2m])) * 100) >=50
         for: 2m
         labels:
           severity: critical
           cce_alert_kind: resources
           alertname:  集群CPU使用率超过50%
           kind: resources
           resource_kind: Cluster
           resourceType: Cluster
           source: prometheus
         annotations:
           info: "集群CPU实际使用率超过50%, 集群当前CPU使用率为{{ printf \"%.2f\" $value }}%"
           description: "集群CPU实际使用率超过50%, 集群当前CPU使用率为{{ printf \"%.2f\" $value }}%"
   ```
   
2. 配置成功后，可以访问Prometheus的Web页面，在"Alert"页面查询告警规则是否触发或者生效。 ![](https://support.huaweicloud.com/usermanual-cce/zh-cn_image_0000001709216590.png "点击放大")
   
3. Prometheus插件将自动推送告警至Alertmanager。如需配置告警接收方，可以通过配置monitoring命名空间下名称为alertmanager的密钥来实现。 ![](https://support.huaweicloud.com/usermanual-cce/zh-cn_image_0000001748700916.png "点击放大")
   ![](https://support.huaweicloud.com/usermanual-cce/public_sys-resources/caution_3.0-zh-cn.png)
   查看alertmanager-alertmanager有状态负载的YAML可以发现，告警数据存放在Pod磁盘中，如果Pod重启，告警数据就会消失。如需持久化，请规划一个PVC，并修改Alertmanager的CR资源，挂载PVC。
   
 
#### 相关文档
- [Prometheus告警规则](https://prometheus.io/docs/prometheus/latest/configuration/alerting_rules/)
- [Prometheus记录规则](https://prometheus.io/docs/prometheus/latest/configuration/recording_rules/)
- [Alertmanager配置说明](https://prometheus.io/docs/alerting/latest/configuration/)
 
