Configuring SLIs
In cloud O&M management scenarios, enterprise core businesses are deployed based on the cloud native architecture. Multiple service instances run collaboratively and multi-dimensional requests interact with each other. O&M metrics are scattered and there is no unified quantification standard and clear baseline for good service quality. As a result, monitoring and warning lack accurate basis, and there is no quantification standard for fault determination after a fault occurs. The technical and business teams have different understanding of service stability, and cross-team collaboration efficiency is low. Based on the quantified data of request- and instance-based SLIs, SLOs are set for each business domain and service level to specify core quality objectives, such as availability, fault rectification duration, and resource load thresholds. This sets service quality red lines, provides clear warning baselines for O&M monitoring, quantified criteria for fault handling, and unified service quality consensus for cross-team collaboration. In addition, data-driven decision-making support is provided for resource allocation and service iteration and optimization, shifting service quality control from passive response to proactive prediction.
Metrics are classified into request- and instance-based SLIs. After the configuration is complete, you can call the metrics in Configuring SLO Interruption Records.
In this module, you can add, modify, delete, and view SLI metrics.
Difference and Relationship Between SLIs and SLOs
Service Level Indicator (SLI): is a specific metric used to objectively and quantitatively measure the key dimensions of a service. It is a quantitative data carrier of service quality. The key dimensions of a service cover the request type (such as request success rate and response time) and instance type (such as instance availability rate and CPU usage). Specific values or ratios are used to reflect the actual operational status of the service. The values or ratios are the data basis for setting the SLO and only objectively present the actual service performance.
Service Level Objective (SLO): is the quantitative expectation and threshold of service quality based on SLIs. It is the red line of service quality. Based on service requirements and O&M capabilities, SLOs define clear target ranges for core SLIs (for example, the request success rate is greater than or equal to 99.95%, and the instance availability rate is greater than or equal to 99.9%). SLOs are the consensus on service quality between technical team and business team. They are also the core criteria for O&M monitoring, troubleshooting, and resource optimization.
| Dimension | SLI | SLO |
|---|---|---|
| Core positioning | Quantitative measurement basis of service quality | Quantitative target standard of service quality |
| Core attribute | Objectivity and data (no goal orientation) | Subjectivity and goal orientation (based on business/O&M requirements) |
| Form | Specific value or ratio (for example, request success rate: 99.98%; CPU usage: 60%) | Value threshold or range (for example, request success rate: ≥ 99.95%; CPU usage: ≤ 80%) |
| Core function | Objectively reflect the actual service operational status. | Define the red line of service quality, guide O&M actions, and align cross-team consensus. |
| Relationship | SLI is the basis and prerequisite for setting the SLO. If there is no SLI, there is no quantitative SLO. | An SLO defines the target threshold for a given SLI. The SLO needs to be verified by monitoring the SLI. |
SLI and SLO are the core concepts of service quality control in cloud O&M. SLOs are set based on SLIs. An SLA is a formal agreement between users and business parties. These three concepts form a service quality control system covering measurement, targets, and agreements.
Prerequisites
You have enabled the fault management package. For details about billing, see Billing Items.
Adding an SLI
- Log in to COC.
- In the navigation pane, choose Basic Configurations > SLO Management.
- Locate the SLO metric you want to configure and click Configure Metric in the Operation column.
- Click Add SLI Metric.
- Set parameters for adding an SLI.
Table 2 Parameters for adding an SLI Parameter
Description
Metric Type
The options are Request SLI metric and Instance SLI metric.
- Request SLI metric: reflects request processing capability and service quality. It is a business- or user-oriented SLI metric. For example, request success rate, response duration, throughput, and error rate.
- Instance SLI metric: reflects the health status, resource usage, and operational stability of service operation carriers. It is an infrastructure- or service deployment-oriented SLI metric. For example, instance availability rate, CPU usage, memory usage, disk I/O load, and port survival status.
SLI Metric Name
Specify an SLI name.
The value can contain 3 to 100 characters, including letters, digits, hyphens (-), and underscores (_).
SLI Metric Description
Describe the SLI. The value contains a maximum of 256 characters.
Definition of Unavailability
Select the parameter comparison method, parameter value, and parameter unit.
The service is unavailable when the parameter meets the conditions.
- Click OK.
- After the setting is complete, click OK.
After the SLI is added, you can view it under the corresponding SLO.
Figure 1 Viewing an SLI
More Operations
You can also perform the following operations.
| Function | Scenario | Operation |
|---|---|---|
| Modifying an SLI Metric | Modify a created SLI based on service requirements. |
|
| Deleting an SLI Metric | Delete an SLI that is no longer applicable. |
|
What is your overall rating for this page?
Thank you very much for your feedback. We will continue working to improve the documentation.See the reply and handling status in My Cloud VOC.
For any further questions, feel free to contact us through the chatbot.
Chatbot