Best Practices of AI Skill-based UCS Multi-Cluster Fleet Policy Management and Compliance
Huawei Cloud UCS is a multi-cluster fleet management service. It provides fleet-level cluster management and classification, and enables consistent compliance governance through policy instances. This practice describes how to use the AI CLI tool to load the huawei-cloud-ucs-policy-governor skill and use a prompt to complete the end-to-end governance process from discovering policy definitions to creating fleet policy instances, enabling policy execution and performing compliance audits.
Prerequisites
- You have installed AI CLI and configured the Huawei Cloud credential environment variables HUAWEI_CLOUD_AK, HUAWEI_CLOUD_SK, and HUAWEI_CLOUD_REGION.
- You have installed hcloud CLI (version 7.2.2 or later). If you use the CLI for the first time, use printf "y\n" | hcloud version to accept the privacy statement.
- CCE clusters or self-managed Kubernetes clusters have been managed by UCS and added to fleets. This practice assumes that you have one or more CCE clusters or self-managed clusters that are already managed by UCS and have been grouped into fleets on the UCS console. If they are not, use the huawei-cloud-ucs-cluster-onboarding-manager skill to manage clusters in UCS. For details, see Creating and Managing a Fleet.
- The huawei-cloud-ucs-policy-governor skill uses the preview-first design. Read operations, such as querying policy definitions and viewing execution status, can be performed immediately. Write operations, such as creating policy instances, enabling policy execution, and disabling policies, first return a preview, and the changes are applied only after user confirmation.
IAM Permission Requirements
| API Action | Permissions | Purpose |
|---|---|---|
| ucs:clusterPolicyInstance:create | Creating a policy | Creating a policy instance for a cluster |
| ucs:clusterGroupPolicyInstance:create | Creating a policy | Creating a fleet policy instance |
| ucs:policyInstance:update | Updating a policy | Modifying a policy instance |
| ucs:policyInstance:get | Obtaining a policy | Viewing policy instance details |
| ucs:policyInstance:delete | Deleting a policy | Deleting a policy instance |
| ucs:policyInstance:list | Listing policies | Listing all policy instances |
| ucs:policyDefinition:list | Listing definitions | Listing available policy definitions |
| ucs:policyDefinition:get | Obtaining a definition | Viewing policy definition details |
| ucs:clusterPolicy:enable | Enabling a policy | Enabling a policy for a cluster |
| ucs:clusterPolicy:disable | Disabling a policy | Disabling a policy for a cluster |
| ucs:clusterGroupPolicy:enable | Enabling a policy | Enabling a policy for a fleet |
| ucs:clusterGroupPolicy:disable | Disabling a policy | Disabling a policy for a fleet |
| ucs:policyJob:list | Listing jobs | Listing policy execution jobs |
| ucs:policyJob:get | Obtaining a task | Viewing policy execution job details |
Procedure
- View policy definitions.
Before creating a policy instance, understand the available policy definitions and find the appropriate constraintTemplateID.
Enter the following prompt in the AI CLI:
My cluster has been managed by UCS. Please help me create a security baseline policy instance for the production fleet. Enable the policy and gradually upgrade the policy from the enforcement action warn to deny. Preview each write operation and execute it only after I confirm it.
The skill first calls ListPolicyDefinitions to return a list of all available policy definitions. Policy definitions are classified into the following:
- Cluster security policies: security baselines, pod security standards, and privileged container restrictions
- Compliance policies: CIS benchmarks, compliance audits, and regulatory compliance
- Resource policies: resource quotas, resource limits, and cost optimization
- Network policies: network policies, inbound/outbound restrictions, and service mesh rules
Select a policy definition based on recommendations. The skill evaluates the cluster management status and fleet membership to determine the policy execution risks.
- Create a policy instance for a fleet.
After confirming the policy definition, the skill calls CreateClusterGroupPolicyInstance to create a policy instance in the fleet. Set enforcementAction to warn. Initially, the warn mode is recommended, where violations are reported but not blocked.
The skill previews the solution, including:
- Policy definition name and namespace
- Target fleet and member clusters
- Enforcement action (warn)
- Policy parameter configuration
The skill performs the creation operation only after the user confirms it.
Use a scope-specific operation to create a policy instance. To create a policy instance for a cluster, use CreateClusterPolicyInstance (parameter: clusterid). To create a policy instance for a fleet, use CreateClusterGroupPolicyInstance (parameter: clustergroupid). There is no general CreatePolicyInstance operation.
- Enable the policy.
After a policy instance is created, enable the policy for the policy to take effect. The skill will call EnableClusterGroupPolicy to automatically enable the policy.
Use a scope-specific operation to enable a policy. To enable a policy for a cluster, use EnableClusterPolicy (parameter: clusterid). To enable a policy for a fleet, use EnableClusterGroupPolicy (parameter: clustergroupid). There is no general EnablePolicy operation.
- View the job status.
After the policy is enabled, the skill calls ListPolicyJobs to view the job status of the policy and calls ShowPolicyJob to view the details of the job.
Statuses:
- Success: The policy has been successfully deployed, and the compliance check has taken effect.
- InProgress: The policy is being deployed, and the compliance check is to be started.
- Failed: The policy failed to be executed. You need to view the job details to locate the fault.
If the job fails, the skill provides diagnosis suggestions, for example, the cluster is not managed, the fleet has no members, or the IAM permission is insufficient.
- Perform compliance audit.
After the policy is successfully executed, perform compliance audit to check the overall compliance status of the fleet. The skill calls ListPolicyInstances to view the status of all policy instances and calls ListPolicyJobs and ShowPolicyJob to view the compliance audit results.
If any violation is detected, the skill provides rectification suggestions:
- View violation details and locate affected Kubernetes resources.
- Obtain the cluster access credential using DownloadFederationKubeconfig and rectify the violation in the cluster.
- Disable the policy and then re-enable it to trigger a new job.
- Confirm that the status of the new job is Success.
- Upgrade the policy from warn to deny.
After the violation is rectified and compliance is verified, upgrade enforcementAction of the policy from warn to deny.
- Use UpdatePolicyInstance to update enforcementAction to deny. (The policy instance ID is required. The skill queries the instance ID through ShowPolicyInstance first.)
- Use DisableClusterPolicy to disable the policy, and then use EnableClusterPolicy--retry=true to re-enable the policy to verify compliance.
- Check the job status again. Ensure that all policy instances are in the Available status and all jobs are in the Success status.
- Configure the policy namespace.
After the policy is created, control the exact behavior of the policy based on the namespace and parameter configurations.
Enter the following prompt in the AI CLI to update the policy namespace:Table 1 Namespace configurations Configuration
Description
Scenario
Specific namespaces
Effective only in specific namespaces
Security baseline policies, applicable only to key namespaces
All namespaces
Effective in all namespaces
Default behavior of general policies, common governance rules
No namespace specified
Determined by targetKind defined in the policy
A policy with targetKind set to Pod take effect for all Pods.
Modify the policy namespace to exclude system namespaces such as kube-system.
ListPolicyInstances and ListPolicyDefinitions do not filter parameters and return all data. To locate a specific policy instance, provide policyinstanceid for the skill to use ShowPolicyInstance to view details and then use UpdatePolicyInstance to update parameter configurations.
Common Issue Diagnosis
If you encounter any issues during policy governance, you can describe the issue in AI CLI. The skill will call the corresponding diagnosis tool to return a structured report.
| Symptom | Diagnosis Prompt |
|---|---|
| The policy execution job is always in the Failed status. | The policy execution job is in the Failed status. Please diagnose. |
| The policy instance failed to be created. | Before creating a policy instance, check the policy definition list and find the correct constraintTemplateID. |
| The cluster is not managed by UCS. | Use UCS to manage the cluster before enabling the policy. |
| The fleet does not have any member clusters. | Add clusters to the fleet first. |
| Insufficient IAM permissions | "IAM denied" is displayed. Create a custom policy on the IAM console and grant permissions. |
| The result of calling ListPolicyInstances is not as expected. | View the policy instance and check the instance ID. |
| The policy instance already exists. | "409 Conflict" is displayed. Check whether the policy instance has been registered. |
| The namespace of the policy instance needs to be changed. | To change the namespace of a policy instance, delete the existing policy instance and create a new one. |
| The policy definition list is empty. | List the policy definitions. |
| The compliance audit result is outdated. | Check the fleet compliance status. |
What is your overall rating for this page?
Thank you very much for your feedback. We will continue working to improve the documentation.See the reply and handling status in My Cloud VOC.
For any further questions, feel free to contact us through the chatbot.
Chatbot