Creating a Standard Cluster 2.0
Scenarios
All operations in this section need to be performed on both the local cloud and collaboration cloud. This section describes how to create a standard cluster 2.0 on the local cloud and collaboration cloud. The created cluster is used to build a resource pool federation.
Prerequisites
- Nodes in the standard cluster 2.0 resource pool are created in your VPC network. You need to prepare a VPC for the local cloud and a VPC for the collaboration cloud. The CIDR blocks of the two VPCs do not overlap and are connected through Direct Connect/VPN and Enterprise Router.
- Ensure that at least two subnets can be created in the VPC to ensure network communication in cross-cloud scenarios.
Procedure
- Log in to the ModelArts management console. In the navigation pane, choose Resource Management > Standard Cluster 2.0. Figure 1 Entry for creating a standard cluster 2.0
- Click Create Standard Cluster 2.0. On the displayed page, configure the parameters as described below. For details about the VPC and its subnets used by the resource pool of the standard cluster 2.0, see Table 2.
Table 1 Cluster flavor parameters Parameter
Description
Pool Name
Enter a name.
NOTE:The name must start with a lowercase letter and cannot end with a hyphen (-). Only lowercase letters, digits, and hyphens (-) are allowed. The length ranges from 4 to 15 characters to ensure that the total length of the resource pool ID does not exceed 48 characters.
Job Type
Select a job type.
The options are DevEnviron, Training Jobs, and Inference Service.
NOTE:The standard cluster 2.0 supports only training jobs.
Advanced Cluster Configuration
- Cluster Specifications: Select Default or Custom. If you select Custom, you can set the cluster scale and enable master node HA.
- Cluster scale: indicates the maximum number of instances that can be managed by a resource pool. Select a value based on your service requirements. Only scale-out is supported after the cluster is created.
- Master node HA: Once HA is enabled, the system creates three control plane nodes for your cluster to ensure reliability. If there are 1,000 or 2,000 nodes in the cluster, HA must be enabled. If HA is disabled, only one master node will be created for your cluster. After a resource pool is created, the HA status of master nodes cannot be changed.
- Service Node Provisioning: Select Automatic or Manual. It is recommended that master nodes be randomly deployed in different AZs to improve DR.
- Automatic: The system randomly allocates AZs for master nodes. Master nodes are randomly distributed in different AZs to improve DR as much as possible. If resources in an AZ are insufficient, master nodes will be distributed in an AZ with sufficient resources to ensure successful cluster creation. In this case, AZ-level DR may not be ensured.
- Manual: Specify AZs for master nodes.
Table 2 Network parameters Parameter
Description
Virtual Private Cloud
Select a VPC for the new cluster.
The VPC you selected must be the same as that selected in step 2. Ensure that the VPCs of the local cloud and collaboration cloud have been connected through Direct Connect/VPN and Enterprise Router.
Default node subnet
Select a subnet. Once selected, all nodes in the resource pool will automatically use the IP addresses assigned within that subnet.
Container subnet
Select the subnet where the container is located. After this parameter is specified, Cloud Native Network 2.0 (ENI network) will be used. The container subnet determines the maximum number of containers in a cluster. This parameter is mandatory for creating a standard cluster 2.0.
Cross-parameter subnet
Used for cross-parameter-plane communication between the local cloud and collaboration cloud. The cross-parameter-plane subnet cannot be the same as the default node subnet or container subnet. This parameter must be specified in cross-cloud scenarios.
K8s CIDR Block
Select Default or Custom.
- Default: The system randomly allocates a CIDR block that does not conflict with other CIDR blocks to you. This CIDR block cannot be changed after creation. Therefore, you are advised to manually configure a CIDR block that best serves your commercial services.
- Custom: You need to specify custom values for K8S Container Network and K8S Service Network.
- K8S Container Network: used by the container in a cluster, which determines how many containers there can be in a cluster. The value cannot be changed after creation.
- K8S Service Network: used when the containers in the same cluster access each other, which determines how many Services there can be. The value cannot be changed after creation.
Service CIDR Block
Enter a proper CIDR block based on the value of K8s CIDR Block.
Table 3 Default specifications Parameter
Description
CPU Architecture
CPU architecture of the resource type. Currently, the CPU supports x86 and Arm architectures, as well as heterogeneous scheduling for x86 and Arm64. Set this parameter as needed.
- x86: If you use GPU resources, select x86. It applies to most general-purpose compute scenarios and supports a wide range of software ecosystems.
- ARM64: If you use NPU resources, select Arm. It applies to specific optimization scenarios such as mobile applications and embedded systems and features low power consumption.
Instance Specifications Type
Select CPU, GPU, or Ascend as needed. Select the CPU architecture and then the instance specifications as needed. The specifications vary depending on the region.
- CPU: a general-purpose compute architecture with low compute performance, suitable for general-purpose tasks.
- GPU: a parallel compute architecture with high compute performance, suitable for parallel tasks. It supports multi-PU distributed training, suitable for scenarios such as deep learning training and image processing.
- Ascend: a dedicated AI architecture with high compute performance, suitable for AI tasks. It supports multi-node distributed deployment, suitable for scenarios such as AI model training and inference acceleration.
Instance Specifications
Select the required specifications from the drop-down list. Due to system loss, the actual available resources are less than those specified in the specifications. After a dedicated resource pool is created, view the available resources on the Nodes tab of the details page.
Contact your account manager to request restricted specifications (such as Ascend) in advance. The specifications will be available within one to three working days. If there is no account manager, submit a service ticket.
AZ
Select Automatic or Manual as needed. An AZ is a physical region where resources use independent power supplies and networks. AZs are physically isolated but interconnected over an intranet. To improve workload reliability, you are advised to create cloud servers in different AZs.
- Automatic: AZs are automatically allocated.
- Manual: Specify AZs for resource pool nodes. To ensure system disaster recovery, deploy all nodes in the same AZ. You can set the number of instances in an AZ.
Instances
Set the number of instances in a dedicated resource pool. More instances mean higher compute performance.NOTE:- If you select Manual for AZ, the number of instances is automatically calculated based on the AZ data. You do not need to set this parameter again.
- Do not create more than 30 instances at a time. Otherwise, the creation may fail due to traffic limiting.
Advanced Specifications Configuration
After enabling this option, you can set the container engine space size and instance OS.
Table 4 Advanced configurations Parameter
Description
Cluster Description
Enter the cluster description to facilitate cluster search and identification.
Tags
Click Add to configure tags for the cluster so that resources can be managed by tag.
The tag information can be predefined in Tag Management Service (TMS) or custom. You can also set tag information on the Tags tab of the details page after the standard cluster 2.0 is created.
Predefined TMS tags are available to all service resources that support tags. Custom tags are available only to the service resources of the user who has created the tags.
Figure 2 Example of the network configuration page
- Cluster Specifications: Select Default or Custom. If you select Custom, you can set the cluster scale and enable master node HA.
- Click Create now.
- After the creation is complete, you can view the created cluster resource pool on the standard cluster 2.0 page.
- Switch from the local cloud to the collaboration cloud. Create a standard cluster 2.0 on the collaboration cloud by taking 2 to 4.
What is your overall rating for this page?
Thank you very much for your feedback. We will continue working to improve the documentation.See the reply and handling status in My Cloud VOC.
For any further questions, feel free to contact us through the chatbot.
Chatbot