HPA
A Horizontal Pod Autoscaler (HPA) is a resource object that dynamically adjusts the number of pods in a Deployment based on certain metrics so that the Deployment can adapt to metric changes. HPA automatically increases or decreases the number of pods based on the configured policy to meet your requirements.
Notes and Constraints
Each pod in CCI 2.0 runs in an environment where system components are installed. These system components occupy some resources, which may cause the resource usage of the pod to be lower than the expected. As a result, the triggering conditions of the HPA policy you configure may be affected. To avoid this, you can reserve system overhead.
Creating and Managing an HPA Policy on the Console
- Log in to the CCI 2.0 console.
- In the navigation pane, choose Workloads. On the Deployments tab, locate the target Deployment and click its name.
- On the Deployment details page, click the Auto Scaling tab.
- Click Create HPA Policy and supplement related information as prompted. For details about the parameters, see Table 1.
Table 1 Parameters for creating an HPA policy Parameter
Description
Name
Enter an HPA policy name.
Namespace
Select a namespace. If you need to create a namespace, click Create Namespace.
Workload
CCI automatically matches the associated workload.
Pod Range
Enter the maximum and minimum numbers of pods that can be scaled by HPA.
Scaling Behavior
- Default
Workloads will be scaled using the Kubernetes default behavior.
- Custom
Workloads will be scaled using custom policies such as the stabilization window, steps, and priorities. Unspecified parameters use the values recommended by Kubernetes.
System Policies
- Metric: the metric type that triggers auto scaling. You can select CPU usage or memory usage.
NOTE:Calculation method: Usage = Current CPU or memory usage by pods /Requested CPUs or memory
- Desired Value: The ideal resource utilization for the workload. This value serves as the baseline for calculating the desired number of pods using the formula: Round up (Current metric value/Desired value × Current number of pods).
Example: Assume that the current number of pods is 2 and the desired value is 50%. If the actual metric value increases to 85%, the system rounds up the value calculated using the formula (85%/50% × 2) = (3.4) to 4. In this case, the number of pods needs to be scaled out to 4.
- Tolerance Range: To prevent flapping (frequent scaling caused by short-term metric fluctuations), a tolerance mechanism is applied. The default tolerance is 0.1 (10%).
- No-scaling range: If the actual metric value falls within [Desired value × (1 – Tolerance), Desired value × (1 + Tolerance)], no scaling action is triggered. The current state is considered acceptable.
- Example: If the desired value is 50% and the tolerance is 0.1, the no-scaling range is 45% to 55%. Scaling calculations are performed only when actual usage is continuously exceeds 55% (scale-out) or falls below 45% (scale-in).
If the metric value sits between the scale-in and scale-out thresholds, no scaling occurs. This parameter is available only in clusters v1.15 or later.
NOTICE:Multiple scaling policies can be configured.

- Default
- Optional: Click Update HPA Policy or Delete HPA Policy to update or delete a created HPA policy.
Creating an HPA Policy Using a YAML File
- If spec.metrics.resource.target.type is set to Utilization, you need to specify the resource requests when creating a workload.
- When spec.metrics.resource.target.type is set to Utilization, the resource usage is calculated as follows: Resource usage = Used resource/Available pod flavor. You can determine the actual flavor of a pod by referring to Pod Flavor.
A properly configured auto scaling policy includes the metric, threshold, and step. It eliminates the need to manually adjust resources in response to service changes and traffic bursts, thus helping you reduce workforce and resource consumption. Currently, CCI supports only one type of auto scaling policy:
- Log in to the CCI 2.0 console.
- In the navigation pane, choose Workloads. On the Deployments tab, locate the target Deployment and click its name.
- On the Auto Scaling tab, click Create from YAML to configure a policy.
The following are example files in different formats:
- Resource description in the hpa.yaml file
kind: HorizontalPodAutoscaler apiVersion: cci/v2 metadata: name: nginx # HPA name namespace: test # HPA namespace spec: scaleTargetRef: # Reference the resource to be automatically scaled. kind: Deployment # Type of the target resource, for example, Deployment name: nginx # Name of the target resource apiVersion: cci/v2 # Version of the target resource minReplicas: 1 # Minimum number of replicas for HPA scaling maxReplicas: 5 # Maximum number of replicas for HPA scaling metrics: - type: Resource # Resource metrics are used. resource: name: memory # Resource name, for example, cpu or memory target: type: Utilization # Metric type. The value can be Utilization (percentage) or AverageValue (absolute value). averageUtilization: 50 # Resource usage. For example, when the CPU usage reaches 50%, scale-out is triggered. behavior: scaleUp: stabilizationWindowSeconds: 30 # Scale-out stabilization duration, in seconds policies: - type: Pods # Number of pods to be scaled value: 1 periodSeconds: 30 # The check is performed once every 30 seconds. scaleDown: stabilizationWindowSeconds: 30 # Scale-in stabilization duration, in seconds policies: - type: Percent # The resource is scaled in or out based on the percentage of existing pods. value: 50 periodSeconds: 30 # The check is performed once every 30 seconds. - Resource description in the hpa.json file
{ "kind": "HorizontalPodAutoscaler", "apiVersion": "cci/v2", "metadata": { "name": "nginx", # HPA name "namespace": "test" # HPA namespace }, "spec": { "scaleTargetRef": { # Reference the resource to be automatically scaled. "kind": "Deployment", # Type of the target resource, for example, Deployment "name": "nginx", # Name of the target resource "apiVersion": "cci/v2" # Version of the target resource }, "minReplicas": 1, # Minimum number of replicas for HPA scaling "maxReplicas": 5, # Maximum number of replicas for HPA scaling "metrics": [ { "type": "Resource", # Resource metrics are used. "resource": { "name": "memory", # Resource name, for example, cpu or memory "target": { "type": "Utilization", # Metric type. The value can be Utilization (percentage) or AverageValue (absolute value). "averageUtilization": 50 # Resource usage. For example, when the CPU usage reaches 50%, scale-out is triggered. } } } ], "behavior": { "scaleUp": { "stabilizationWindowSeconds": 30, "policies": [ { "type": "Pods", "value": 1, "periodSeconds": 30 } ] }, "scaleDown": { "stabilizationWindowSeconds": 30, "policies": [ { "type": "Percent", "value": 50, "periodSeconds": 30 } ] } } } }
- Resource description in the hpa.yaml file
- Click OK. You can view the policy on the Auto Scaling tab.Figure 1 Auto scaling policy
When the trigger condition is met, the auto scaling policy will be executed.
Feedback
Was this page helpful?
Provide feedbackThank you very much for your feedback. We will continue working to improve the documentation.See the reply and handling status in My Cloud VOC.
For any further questions, feel free to contact us through the chatbot.
Chatbot