Updated on 2026-09-08 GMT+08:00

Scheduling Policy

Overview

In Huawei Cloud resource management scenarios, users configuring inference services may need to optimize scheduling policies to enhance service performance or resource utilization. For instance, to ensure high availability, users typically select a high-availability scheduling policy, which requires service replicas (pods) to be distributed as evenly as possible across different nodes to minimize the impact of individual failures. However, in practice, users may find that existing scheduling configuration options lack the flexibility required for specific business needs, such as scenarios needing finer control over scheduling priorities, or those requiring a compact scheduling policy to maximize resource utilization. Faced with this challenge, users often wonder: How can scheduling policies be configured more flexibly on Huawei Cloud to adapt to diverse service requirements?

To address this, ModelArts offers three distinct scheduling policies for real-time inference services in dedicated resource pools: high-availability scheduling, compact scheduling, and affinity scheduling. By flexibly adjusting the distribution logic of service instances across cluster nodes, these policies satisfy three core service demands: highly reliable operations, efficient resource utilization, and custom deployment control.

Supported Scheduling Policies

Inference services support the coordination of three scheduling policies: high-availability scheduling, compact scheduling, and affinity scheduling.

  • High-availability scheduling: Requires service replicas (pods) to be distributed as evenly as possible across different nodes to minimize the impact of individual failures.
  • Compact scheduling: Also known as the binpack scheduling policy. It reduces resource fragmentation to maximize utilization and takes effect across the entire cluster.
  • Affinity scheduling: node affinity and anti-affinity.
Table 1 Comparison of scheduling policies

Dimension

High-Availability Scheduling

Compact Scheduling

Affinity Scheduling

Core purpose

Resists individual failures; guarantees stable and continuous operations.

Consolidates idle resources; improves overall resource utilization.

Customizes deployment scope; achieves dedicated service control and isolation.

Use case

Online core production business; real-time, high-concurrency inference.

Testing and debugging; gray services; low-traffic, non-core services.

Model warmup deployment; hardware binding; service security isolation.

Instance distribution

Dispersed and distributed across multiple different physical nodes.

Concentrated and gathered into the minimum number of nodes.

Target-deployed or target-avoided based on custom rules.

Compatible resource pools

Dedicated resource pools only.

Dedicated resource pools only.

Dedicated resource pools only.

Failure risk

Dispersed risk; individual failures have a minimal impact.

Concentrated risk; node failure can easily batch affect services.

Depends on the status of selected nodes; strong rules offer low fault tolerance.

Resource utilization

Moderately stable.

Highest and optimal.

Controllable on demand.

Core constraints

Number of replicas ≥ 2; relies on sufficient cluster nodes.

Not for core production; limited to node pools of the same specification.

Strong rules can easily cause deployment failure; limited number of nodes can be selected.

Basic prerequisites

Dedicated resource pool runs normally; replica count meets the standard.

No strict high-availability requirements; uniform node specifications.

Target nodes are healthy and available; warmup/node partitioning completed in advance.

Quick configuration points

Raise scheduling priority; configure rolling upgrades and self-healing.

Turn off affinity scheduling; lower scheduling priority.

Enable scheduling; select type and strength; specify nodes.

Among the three scheduling policies mentioned above, when node affinity and node anti-affinity are configured as strong affinity, they act as strong constraints that the scheduler must satisfy unconditionally. All other scenarios are treated as weak constraints, which the scheduler will make every effort to satisfy first. However, in cases where node resources are extremely tight or the physical resource pool is large (for example, when the number of nodes exceeds 100), there is a possibility that these constraints cannot be fully met.

High-Availability Scheduling

Concepts

High-availability scheduling is the default recommended policy for real-time inference services. Its core objective is to eliminate individual failure and guarantee continuous service availability. This policy is ideal for core production scenarios with zero-downtime tolerances and stringent stability requirements, such as online intelligent customer service, real-time financial risk control, and autonomous driving decision-making inference. By distributing service replicas across different physical nodes, racks, or even data centers, it prevents a single-node or single-rack failure from causing total service unavailability. Furthermore, it supports seamless transitions during rolling upgrades, ensuring zero service interruption.

Configuration entry

Log in to the ModelArts console and choose Model Inference > Real-Time Inference. When deploying a real-time service, select HA scheduling in Resource Settings > Scheduling Policy. For details about how to deploy a real-time service, see Configuring Deployment Settings.

Constraints

  • You can select either HA scheduling or Compact scheduling. By default, HA scheduling is enabled.
  • If both HA scheduling and Affinity Scheduling are enabled, Affinity Scheduling takes precedence.

Compact Scheduling

Concepts

The core objective of compact scheduling is to maximize resource utilization and reduce deployment costs. Enable bin packing for cluster workloads. The scheduler will prioritize placing pods on nodes with higher resource consumption to reduce idle resource fragmentation and improve cluster resource utilization.

Use cases

This mode works well for situations needing moderate stability and lower costs, like non-core tests, offline inference transitions, and edge deployments with few resources. Common uses are checking model functions, low-traffic tests, and light inference on edge devices.

Configuration entry

Log in to the ModelArts console and choose Model Inference > Real-Time Inference. When deploying a real-time service, select Compact scheduling in Resource Settings > Scheduling Policy. For details about how to deploy a real-time service, see Configuring Deployment Settings.

Constraints

  • You can select either HA scheduling or Compact scheduling. By default, HA scheduling is enabled.
  • When both Compact scheduling and Affinity Scheduling are enabled, Affinity Scheduling takes precedence.

Affinity Scheduling

Concepts

Affinity scheduling allows you to control where service replicas run by using node affinity or anti-affinity rules. This helps meet specific deployment needs. You can configure the node affinity type and strength to flexibly schedule workloads in a resource pool. If no nodes are specified, the pods will be randomly scheduled according to the default cluster scheduling policy.

Configuration entry

Log in to the ModelArts console and choose Model Inference > Real-Time Inference. When deploying a real-time service, select Affinity Scheduling in Unit Settings > More Settings. In the dialog box displayed on the right, configure the affinity scheduling policy. For details about how to deploy a real-time service, see Configuring Deployment Settings.

Table 2 Affinity scheduling configuration

Affinity Type

Strength

Use Case

Node affinity

Weak affinity

Prioritizes deployment to specified nodes. If node resources are insufficient, scheduling can fall back to other nodes, balancing cache reuse with resource flexibility.

Strong affinity

Enforces deployment to specified nodes. Ideal for dedicated nodes with pre-warmed models, nodes with local hardware encryption cards mounted, or edge nodes bound to specific peripherals (such as cameras or sensors), ensuring model cache reuse and exclusive hardware resource utilization.

Node anti-affinity

Weak affinity

Avoids deployment to specified nodes as much as possible, but allows compromised deployment if no other nodes are available, balancing isolation requirements with resource availability.

Strong affinity

Prohibits deployment to specified nodes. Ideal for preventing co-location with high-load services, avoiding fault-prone nodes, and isolating sensitive services from non-sensitive services.

In the Add Node list, select the nodes that meet the preceding configuration rules.

In the affinity scheduling policy, a single account (not an IAM user) can add a maximum of 20 nodes by default. To increase the number of API keys that can be created, submit a service ticket.

When you choose a pre-warmed model, only the pre-warmed nodes appear on the affinity scheduling page. Unwarmed nodes do not show. A message appears stating: The selected model is pre-warmed and will deploy automatically on the best node. Specifying a node might cause warmup to fail.