Help Center/ MapReduce Service/ API Reference/ API V2/ Cluster Management APIs/ Creating a Cluster and Submitting a Job - CreateClusterAndSubmitJob
Updated on 2026-09-15 GMT+08:00

Creating a Cluster and Submitting a Job - CreateClusterAndSubmitJob

Function

This API is used to create an MRS cluster, submit a job, and terminate the cluster after the job is complete. This API is supported in MRS 1.8.9 or later. Before using this API, you need to obtain the following resource information:

  • Create or query a VPC and subnet.
  • Create or query a key pair using an ECS.
  • Obtain the region information by referring to Endpoints.
  • Obtain the MRS version and the components supported by the MRS version by referring to Obtaining the MRS Cluster Information.

Constraints

None

Debugging

You can debug this API through automatic authentication in API Explorer. API Explorer can automatically generate sample SDK code and supports sample SDK code debugging.

Authorization Information

Each account has all the permissions required to call all APIs, but IAM users must be assigned the required permissions.

  • If you are using role/policy-based authorization, see Permissions Policies and Supported Actions for details on the required permissions.
  • If you are using identity policy-based authorization, the following identity policy-based permissions are required.

    Action

    Access Level

    Resource Type (*: required)

    Condition Key

    Alias

    Dependency

    mrs:cluster:createCluster

    Write

    cluster *

    -

    • mrs:cluster:create
    • iam:agencies:pass
    • ecs:cloudServers:createServers
    • ecs:cloudServers:deleteServers
    • ecs:cloudServerQuotas:get
    • ecs:cloudServerFlavors:get
    • ecs:cloudServers:listServersDetails
    • ecs:cloudServers:showServerGroup
    • ecs:cloudServers:updateMetadata
    • ecs:cloudServers:start
    • ecs:cloudServers:stop
    • ecs:serverGroups:manage
    • ecs:cloudServers:listServerInterfaces
    • vpc:vpcs:list
    • vpc:vpcs:create
    • eip:publicIps:list
    • eip:publicIps:get
    • vpc:ports:get
    • vpc:ports:create
    • vpc:ports:delete
    • vpc:ports:update
    • vpc:privateIps:create
    • vpc:privateIps:delete
    • vpc:securityGroups:get
    • vpc:securityGroups:create
    • vpc:securityGroups:delete
    • vpc:securityGroupRules:create
    • vpc:securityGroupRules:delete
    • vpc:quotas:list
    • evs:quotas:get
    • evs:types:get
    • rds:instance:get
    • rds:instance:listAll
    • kms:cmk:list
    • vpc:floatingIps:get
    • bms:serverQuotas:get
    • bms:serverFlavors:get
    • bms:servers:start
    • bms:servers:stop
    • bms:servers:create
    • bms:servers:updateMetadata
    • bms:servers:showBaremetalServer
    • bms:servers:list

    -

    • g:RequestTag/<tag-key>

    • g:TagKeys

    • g:EnterpriseProjectId

URI

POST /v2/{project_id}/run-job-flow

Table 1 URI parameters

Parameter

Mandatory

Type

Description

project_id

Yes

String

Definition

Project ID. For details about how to obtain the project ID, see Obtaining a Project ID.

Constraints

N/A

Range

The value must consist of 1 to 64 characters. Only letters and digits are allowed.

Default Value

N/A

Request Parameters

Table 2 Request body parameters

Parameter

Mandatory

Type

Description

is_dec_project

No

Boolean

Definition

Whether the resource is a DeC resource, that is, whether the cluster is a DeC cluster.

Constraints

N/A

Range

  • true: The resource is a DeC resource.

  • false: The resource is not a DeC resource.

Default Value

false

cluster_version

Yes

String

Definition

Cluster version, for example, MRS 3.1.0.

Constraints

N/A

Range

N/A

Default Value

N/A

cluster_name

Yes

String

Definition

Cluster name.

Constraints

N/A

Range

The cluster name must globally unique.

A cluster name can contain only 1 to 64 characters. Only letters, numbers, hyphens (-), and underscores (_) are allowed.

Default Value

N/A

cluster_type

Yes

String

Definition

The cluster type.

Constraints

N/A

Range

  • ANALYSIS: analysis cluster
  • STREAMING: streaming cluster
  • MIXED: hybrid cluster
  • CUSTOM: custom cluster, which is supported only by MRS 3.x.

Default Value

N/A

charge_info

No

ChargeInfo object

Definition

The billing type. For details, see Table 7.

Constraints

N/A

Range

N/A

Default Value

N/A

region

Yes

String

Definition

Information about the region where the cluster is located. For details, see Endpoints.

Constraints

N/A

Range

N/A

Default Value

N/A

vpc_name

Yes

String

Definition

The name of the VPC where the subnet is located. Obtain the VPC name by performing the following operations on the VPC management console:

  1. Log in to the VPC management console.
  2. Choose Virtual Private Cloud > My VPCs. On the Virtual Private Cloud page, obtain the VPC name from the list.

Constraints

N/A

Range

N/A

Default Value

N/A

subnet_id

No

String

Definition

The subnet ID. Obtain the subnet ID by performing the following operations on the VPC management console:

  1. Log in to the VPC management console.
  2. Choose Virtual Private Cloud > My VPCs.
  3. Locate the row containing the target VPC and click the number in the Subnets column to view the subnet information.
  4. Click the subnet name to obtain the network ID.

Constraints

At least one of subnet_id and subnet_name must be configured. If the two parameters are configured but do not match the same subnet, the cluster fails to create. subnet_id is recommended.

Range

N/A

Default Value

N/A

subnet_name

Yes

String

Definition

The subnet name. Obtain the subnet name by performing the following operations on the VPC management console:

  1. Log in to the management console.
  2. Choose Virtual Private Cloud > My VPCs.
  3. Locate the row that contains the target VPC and click the number in the Subnets column to obtain the subnet name.

Constraints

At least one of subnet_id and subnet_name must be configured. If the two parameters are configured but do not match the same subnet, the cluster fails to create. If only subnet_name is configured and subnets with the same name exist in the VPC, the first subnet name in the VPC is used when a cluster is created. subnet_id is recommended.

Range

N/A

Default Value

N/A

components

Yes

String

Definition

List of component names, which are separated by commas (,). For details about the components that are supported, see "Components Supported by MRS" in Obtaining the MRS Cluster Information.

Constraints

N/A

Range

N/A

Default Value

N/A

external_datasources

No

Array of ClusterDataConnectorMap objects

Definition

When deploying components such as Hive and Ranger, you can associate data connections and store metadata in associated databases. For details about the parameters, see Table 3.

Constraints

N/A

Range

N/A

Default Value

N/A

availability_zone

Yes

String

Definition

The AZ name. Multi-AZ clusters are not supported. For details about AZs, see Endpoints.

Constraints

N/A

Range

N/A

Default Value

N/A

security_groups_id

No

String

Definition

Security group ID of the cluster. You can view the ID of the security group to be used in the security group list on the VPC management console, or you can create one automatically.

  • If this ID is empty, MRS automatically creates a security group in the background. The name of the automatically created security group starts with mrs_{cluster_name}.

  • If this ID is not empty, a fixed security group is used to create the cluster. The ID passed must be the ID of a security group that has been created in the current tenant.

  • Multiple security group IDs are supported, separated by commas (,).

Constraints

N/A

Range

N/A

Default Value

N/A

auto_create_default_security_group

No

Boolean

Definition

Whether to create the default security group for the MRS cluster.

Constraints

If this parameter is set to true, the default security group will be created for the cluster regardless of whether security_groups_id is specified.

Range

  • true: The default security group is created for the MRS cluster.
  • false: The default security group is not created.

Default Value

false

safe_mode

Yes

String

Definition

The running mode of an MRS cluster.

Constraints

N/A

Range

  • SIMPLE: normal cluster. In a normal cluster, Kerberos authentication is disabled, and users can use all functions provided by the cluster.
  • KERBEROS: security cluster. In a security cluster, Kerberos authentication is enabled, and common users cannot use the file management and job management functions of an MRS cluster or view cluster resource usage and the job records of Hadoop and Spark. To use more functions, the users must obtain the relevant permissions from the Manager administrator.

Default Value

N/A

manager_admin_password

Yes

String

Definition

Password of the MRS Manager administrator.

Constraints

N/A

Range

  • The value must contain 8 to 26 characters.
  • The value must contain at least four of the following: uppercase letters, lowercase letters, numbers, and special characters (!@$%^-_=+[{}]:,./?), but must not contain spaces.
  • The value cannot be the username or the username spelled backwards.

Default Value

N/A

login_mode

Yes

String

Definition

Node login mode.

Constraints

N/A

Range

  • PASSWORD: password-based login. If this value is selected, node_root_password cannot be left blank.
  • KEYPAIR: key pair used for login. If this value is selected, node_keypair_name cannot be left blank.

Default Value

N/A

node_root_password

No

String

Definition

The password of user root for logging in to a cluster node.

Constraints

N/A

Range

  • Must be 8 to 26 characters long.
  • Must contain at least four of the following: uppercase letters, lowercase letters, numbers, and special characters (!@$%^-_=+[{}]:,./?), but must not contain spaces.
  • Cannot be the username or the username spelled backwards.

Default Value

N/A

node_keypair_name

No

String

Definition

The name of a key pair. You can use a key pair to log in to a cluster node.

Constraints

N/A

Range

N/A

Default Value

N/A

enterprise_project_id

No

String

Definition

Enterprise project ID. When you create a cluster, associate the enterprise project ID with the cluster. The default value is 0, indicating the default enterprise project. To obtain the enterprise project ID, see the id value in the enterprise_project field data structure table in "Querying the Enterprise Project List" in Enterprise Management API Reference.

Constraints

N/A

Range

N/A

Default Value

The default value is 0, indicating the default enterprise project.

eip_address

No

String

Definition

EIP bound to an MRS cluster, which can be used to access MRS Manager. The EIP must have been created and must be in the same region as the cluster.

Constraints

N/A

Range

N/A

Default Value

N/A

eip_id

No

String

Definition

ID of the bound EIP.

Constraints

ID of the bound EIP. This parameter is mandatory when eip_address is configured. To obtain the EIP ID, log in to the VPC console, choose Network > Elastic IP and Bandwidth > Elastic IP, click the EIP to be bound, and obtain the ID in the Basic Information area.

Range

N/A

Default Value

N/A

mrs_ecs_default_agency

No

String

Definition

Name of the agency bound to a cluster node by default. The value is fixed to MRS_ECS_DEFAULT_AGENCY. An agency allows ECS or BMS to manage MRS resources. You can configure an agency of the ECS type to automatically obtain the AK/SK to access OBS. The MRS_ECS_DEFAULT_AGENCY agency has the OBS OperateAccess permission of OBS and the CES FullAccess (for users who have enabled fine-grained policies), CES Administrator, and KMS Administrator permissions in the region where the cluster is located.

Constraints

N/A

Range

N/A

Default Value

N/A

template_id

No

String

Definition

The template used for node deployment when the cluster type is CUSTOM.

  • mgmt_control_combined_v2: template for jointly deploying management and controller nodes. The management and controller roles are co-deployed on the master node, and data instances are deployed in the same node group. This deployment model applies to scenarios where there are fewer than 100 nodes, reducing costs.
  • mgmt_control_separated_v2: The management and control roles are deployed on different master nodes, and data instances are deployed in the same node group. This deployment model applies to a cluster with 100 to 500 nodes and delivers better performance in high-concurrency load scenarios.
  • mgmt_control_data_separated_v2: The management and control roles are deployed on different master nodes, and data instances are deployed in different node groups. This deployment model applies to a cluster with more than 500 nodes. Components can be deployed separately, which can be used for a larger cluster scale.

Constraints

N/A

Range

N/A

Default Value

N/A

tags

No

Array of Tag objects

Definition

Cluster tag information. For details, see Table 4.

Constraints

A cluster allows a maximum of 10 tags. A tag name (key) must be unique in a cluster.

Range

N/A

Default Value

N/A

log_collection

No

Integer

Definition

Whether to collect logs when cluster creation fails.

Constraints

N/A

Range

  • 0: Do not create an OBS bucket only for log collection when a cluster fails to be created.
  • 1: Create an OBS bucket only for collect logs when a cluster fails to be created.

Default Value

1

node_groups

Yes

Array of NodeGroupV2 objects

Definition

Information about the node groups that form the cluster. For details about the parameters, see Table 5.

Constraints

N/A

Range

N/A

Default Value

N/A

bootstrap_scripts

No

Array of BootstrapScript objects

Definition

The bootstrap action script. For details about the parameters, see Table 13.

Constraints

N/A

Range

N/A

Default Value

N/A

log_uri

No

String

Definition

The OBS path to which cluster logs are dumped. After the log dump function is enabled, the read and write permissions on the OBS path are required for uploading logs. Configure the default agency MRS_ECS_DEFAULT_AGENCY or customize an agency with the read and write permissions on the OBS path. For details, see Configuring a Storage-Compute Decoupled Cluster (Agency). This parameter is available only for cluster versions that support dumping cluster logs to OBS.

Constraints

N/A

Range

N/A

Default Value

N/A

component_configs

No

Array of ComponentConfig objects

Definition

The custom configuration of cluster components. This parameter applies only to cluster versions that support the feature of creating a cluster by customizing component configurations. For details about this parameter, see Table 14.

Constraints

The number of records cannot exceed 50.

Range

N/A

Default Value

N/A

delete_when_no_steps

No

Boolean

Definition

Whether to automatically delete the cluster after the job is complete.

Constraints

N/A

Range

  • true: The cluster is deleted after the job is complete.
  • false: The cluster is not deleted after the job is complete.

Default Value

false

steps

Yes

Array of StepConfig objects

Definition

The job list. For details about this parameter, see Table 16.

Constraints

The number of records cannot exceed 255.

Range

N/A

Default Value

N/A

Table 3 ClusterDataConnectorMap

Parameter

Mandatory

Type

Description

map_id

No

Integer

Definition

Data connection association ID

Constraints

N/A

Range

N/A

Default Value

N/A

connector_id

No

String

Definition

Data connection ID

Constraints

N/A

Range

N/A

Default Value

N/A

component_name

No

String

Definition

Component name

Constraints

N/A

Range

N/A

Default Value

N/A

role_type

No

String

Definition

Component role type.

Constraints

N/A

Range

  • hive_metastore: Hive Metastore role
  • hive_data: Hive role
  • hbase_data: HBase role.
  • ranger_data: Ranger role

Default Value

N/A

source_type

No

String

Definition

Data connection type

Constraints

N/A

Range

  • LOCAL_DB: local metadata
  • RDS_POSTGRES: RDS PostgreSQL database
  • RDS_MYSQL: RDS MySQL database
  • gaussdb-mysql: TaurusDB

Default Value

N/A

cluster_id

No

String

Definition

ID of the associated cluster

Constraints

N/A

Range

The value can contain 1 to 64 characters, including only letters, digits, underscores (_), and hyphens (-).

Default Value

N/A

status

No

Integer

Definition

Data connection status.

Constraints

N/A

Range

  • 0: Normal.
  • 1: In use.

Default Value

N/A

Table 4 Tag

Parameter

Mandatory

Type

Description

key

Yes

String

Definition

Tag key.

Constraints

N/A

Range

  • A tag key can contain letters, digits, spaces, and special characters _.:=+-@, but cannot start or end with a space or start with _sys_.
  • The tag key of a resource must be unique.
  • It can contain a maximum of 128 Unicode characters and cannot be an empty string.

Default Value

N/A

value

Yes

String

Definition

Tag value.

Constraints

N/A

Range

  • The value can contain letters, digits, spaces, and special characters _.:=+-@, but cannot start or end with a space or start with _sys_.
  • The value can contain a maximum of 255 Unicode characters and can be an empty string.

Default Value

N/A

Table 5 NodeGroupV2

Parameter

Mandatory

Type

Description

group_name

Yes

String

Definition

Node group name.

Constraints

N/A

Range

The value can contain a maximum of 64 characters, including uppercase and lowercase letters, digits and underscores (_). The rules for configuring node groups are as follows:

  • master_node_default_group: master node group, which must be included in all cluster types.
  • core_node_analysis_group: analysis core node group, which must be included in both analysis and hybrid clusters.
  • core_node_streaming_group: streaming core node group, which must be included in both streaming and hybrid clusters.
  • task_node_analysis_group: analysis task node group, which can be selected for analysis clusters and hybrid clusters as needed.
  • task_node_streaming_group: streaming task node group, which can be selected for streaming clusters and hybrid clusters as needed.
  • node_group{x}: node group of a custom cluster. A maximum of nine such node groups can be added for a custom cluster.

Default Value

N/A

node_num

Yes

Integer

Definition

Number of nodes.

Constraints

The total number of Core and Task nodes cannot exceed 500.

Range

0-500

Default Value

N/A

node_size

Yes

String

Definition

Instance specification of the node. For example: {ECS_FLAVOR_NAME}.linux.bigdata, where {ECS_FLAVOR_NAME} can be c3.4xlarge.4 or other ECS specifications visible on the MRS purchase page. For detailed information about instance specifications, see ECS Specifications Used by MRS and BMS Specifications Used by MRS. You are advised to obtain the supported specifications for the corresponding region and version from the cluster creation page of the MRS console.

Constraints

N/A

Range

N/A

Default Value

N/A

root_volume

No

Volume object

Definition

The system disk information of the node. This parameter is optional for some VMs or the system disk of the BMS and mandatory in other cases. For details about this parameter, see Table 6.

Constraints

N/A

Range

N/A

Default Value

N/A

data_volume

No

Volume object

Definition

Data disk information. For details about the parameter, see Table 6.

Constraints

This parameter is mandatory when data_volume_count is not 0.

Range

N/A

Default Value

N/A

data_volume_count

No

Integer

Definition

Number of data disks of a node.

Constraints

N/A

Range

0-20

Default Value

N/A

charge_info

No

ChargeInfo object

Definition

The billing type of a node group. The billing types of master and core node groups are the same as those of the cluster. The billing type of the task node group can be different. For details about this parameter, see Table 7.

Constraints

N/A

Range

N/A

Default Value

N/A

auto_scaling_policy

No

AutoScalingPolicy object

Definition

The auto scaling rule information. For details about this parameter, see Table 8.

Constraints

N/A

Range

N/A

Default Value

N/A

assigned_roles

No

Array of strings

Definition

This parameter is mandatory when the cluster type is CUSTOM. You can specify the roles deployed in a node group. This parameter is a string array. Each string represents a role expression. Role expression definition:

  • If the role is deployed on all nodes in a node group, set this parameter to {role name}, for example, DataNode.
  • If the role is deployed on a specified subscript node in the node group: {role name}:{index1},{index2}…,{indexN}, for example, NameNode:1,2. The subscript starts from 1.
  • Some roles support multi-instance deployment (that is, multiple instances of the same role are deployed on a node): {role name}[{instance count}], for example, EsNode[9]. For details about the available roles, see Roles and components supported by MRS.

Constraints

N/A

Range

N/A

Default Value

N/A

Table 6 Volume

Parameter

Mandatory

Type

Description

type

Yes

String

Definition

Disk type.

Constraints

N/A

Range

  • SATA: common I/O disk
  • SAS: high I/O disk
  • SSD: ultra-high I/O disk
  • GPSSD: general-purpose SSD disk

Default Value

N/A

size

Yes

Integer

Definition

Data disk size in GB.

Constraints

N/A

Range

10-32768

Default Value

N/A

Table 7 ChargeInfo

Parameter

Mandatory

Type

Description

charge_mode

Yes

String

Definition

Billing mode.

Constraints

N/A

Range

  • prePaid: the yearly/monthly billing mode. This mode is now supported for the API used to create a cluster, but is not supported for the API used to create a cluster and submit a job.
  • postPaid: the pay-per-use billing mode.

Default Value

N/A

period_type

No

String

Definition

Period type.

Constraints

N/A

Range

  • month: indicates that the service is charged by month.
  • year: indicates that the fee is charged by year.
  • day: The cluster is billed on a pay-per-use basis.

Default Value

N/A

period_num

No

Integer

Definition

Number of periods.

Constraints

This parameter is valid and mandatory only when charge_mode is set to prePaid.

Range

  • If period_type is set to month, the value ranges from 1 to 9.
  • If period_type is set to year, the value ranges from 1 to 3.

Default Value

N/A

is_auto_pay

No

Boolean

Definition

Whether the order will be automatically paid. This parameter is available for yearly/monthly mode. By default, the automatic payment is disabled.

Constraints

N/A

Range

  • true: The system automatically selects available discounts and coupons, and then pays for the order with the account balances. If the automatic payment fails, an order in Pending payment state is generated waiting for manual payment.
  • false: The user needs to pay for the bill after using available discounts and coupons.

Default Value

false

Table 8 AutoScalingPolicy

Parameter

Mandatory

Type

Description

auto_scaling_enable

Yes

Boolean

Definition

Whether to enable the autoscaling rule.

Constraints

N/A

Range

  • true: Enable the autoscaling rule.

  • false: Disable the autoscaling rule.

Default Value

N/A

min_capacity

Yes

Integer

Definition

Minimum number of nodes reserved for the node group.

Constraints

N/A

Range

0-500

Default Value

N/A

max_capacity

Yes

Integer

Definition

Maximum number of nodes in the node group.

Constraints

N/A

Range

0-500

Default Value

N/A

resources_plans

No

Array of ResourcesPlan objects

Definition

Resource plan list. If this parameter is left blank, resource plans are disabled.

Constraints

When autoscaling is enabled, at least one of resource plans or autoscaling rules must be configured. A maximum of five resource plans are allowed.

Range

N/A

Default Value

N/A

rules

No

Array of Rule objects

Definition

Autoscaling rule list.

Constraints

When autoscaling is enabled, at least one of resource plans or autoscaling rules must be configured. A maximum of 10 rules are allowed.

Range

N/A

Default Value

N/A

exec_scripts

No

Array of ScaleScript objects

Definition

List of custom automation scripts for autoscaling. If this parameter is left blank, automation scripts are disabled. This parameter is currently not supported in the V2 autoscaling policy creation and update API.

Constraints

A maximum of 10 rules are allowed.

Range

N/A

Default Value

N/A

Table 9 ResourcesPlan

Parameter

Mandatory

Type

Description

period_type

Yes

String

Definition

Cycle type of a resource plan. This parameter can be set to daily only.

Constraints

N/A

Range

daily: Charges are calculated by day.

Default Value

N/A

start_time

Yes

String

Definition

Start time of a resource plan. The value is in the format of hour:minute, indicating that the time ranges from 00:00 to 23:59.

Constraints

N/A

Range

N/A

Default Value

N/A

end_time

Yes

String

Definition

End time of a resource plan. The format is the same as that of start_time.

Constraints

The value cannot be earlier than the start_time, and the interval between start_time and start_time cannot be less than 30 minutes.

Range

N/A

Default Value

N/A

min_capacity

Yes

Integer

Definition

Minimum number of the preserved nodes in a node group in a resource plan.

Constraints

N/A

Range

0-500

Default Value

N/A

max_capacity

Yes

Integer

Definition

Maximum number of the preserved nodes in a node group in a resource plan.

Constraints

N/A

Range

0-500

Default Value

N/A

effective_days

No

Array of strings

Definition

The effective date of a resource plan. If this parameter is left blank, it indicates that the resource plan takes effect every day. The options are as follows:

MONDAY, TUESDAY, WEDNESDAY, THURSDAY, FRIDAY, SATURDAY, and SUNDAY

Constraints

N/A

Range

N/A

Default Value

N/A

Table 10 Rule

Parameter

Mandatory

Type

Description

name

Yes

String

Definition

Name of an auto scaling rule.

Constraints

N/A

Range

The value can contain 1 to 64 characters, including only letters, digits, underscores (_), and hyphens (-).

Rule names must be unique in a node group.

Default Value

N/A

description

No

String

Definition

Description about an auto scaling rule.

Constraints

N/A

Range

The value can contain 0 to 1024 characters.

Default Value

N/A

adjustment_type

Yes

String

Definition

Adjustment type of an auto scaling rule.

Constraints

N/A

Range

  • scale_out: cluster scale-out
  • scale_in: cluster scale-in

Default Value

N/A

cool_down_minutes

Yes

Integer

Definition

The cluster cooling time after an auto scaling rule is triggered, in minutes, during which period no auto scaling operation is performed.

Constraints

N/A

Range

The value ranges from 0 to 10080. 10080 indicates the number of minutes in a week.

Default Value

N/A

scaling_adjustment

Yes

Integer

Definition

Number of nodes that can be adjusted once.

Constraints

N/A

Range

1-100

Default Value

N/A

trigger

Yes

Trigger object

Definition

Condition for triggering a rule. For details about this parameter, see Table 11.

Constraints

N/A

Range

N/A

Default Value

N/A

Table 11 Trigger

Parameter

Mandatory

Type

Description

metric_name

Yes

String

Definition

Metric name. This triggering condition makes a judgment according to the value of the metric.

Constraints

N/A

Range

The value can contain 0 to 64 characters.

Default Value

N/A

metric_value

Yes

String

Definition

Metric threshold to trigger a rule. The value must be an integer or a number with two decimal places.

Constraints

N/A

Range

Only integers or numbers with two decimal places are allowed.

Default Value

N/A

comparison_operator

No

String

Definition

Metric judgment logic operator.

Constraints

N/A

Range

  • LT: less than
  • GT: greater than
  • LTOE: less than or equal to
  • GTOE: greater than or equal to

Default Value

N/A

evaluation_periods

Yes

Integer

Definition

Number of consecutive five-minute periods, during which a metric threshold is reached

Constraints

N/A

Range

1-288

Default Value

N/A

Table 12 ScaleScript

Parameter

Mandatory

Type

Description

name

Yes

String

Definition

Names of custom scaling automation scripts.

Constraints

N/A

Range

The names in the same cluster must be unique. The value can contain only numbers, letters, spaces, hyphens (-), and underscores (_) and cannot start with a space. The value can contain 1 to 64 characters.

Default Value

N/A

uri

Yes

String

Definition

Path of a custom automation script. Set this parameter to an OBS bucket path or a local VM path.

  • OBS bucket path: Enter a script path, for example, obs://XXX/scale.sh.
  • Local VM path: Enter a script path. The script path must start with a slash (/) and end with .sh.

Constraints

N/A

Range

N/A

Default Value

N/A

parameters

No

String

Definition

Parameters of a custom automation script. Multiple parameters are separated by space. The following predefined system parameters can be transferred:

  • ${mrs_scale_node_num}: The number of nodes to be added or removed
  • ${mrs_scale_type}: The scaling type. The value can be scale_out or scale_in.
  • ${mrs_scale_node_hostnames}: Host names of the nodes to be added or removed
  • ${mrs_scale_node_ips}: IP addresses of the nodes to be added or removed
  • ${mrs_scale_rule_name}: Name of the rule that triggers auto scaling Other user-defined parameters are used in the same way as those of common shell scripts. Parameters are separated by space.

Constraints

N/A

Range

N/A

Default Value

N/A

nodes

Yes

Array of strings

Definition

Name of the node group where the custom automation script is executed.

Constraints

N/A

Range

N/A

Default Value

N/A

active_master

No

Boolean

Definition

Whether the custom automation script runs only on the active Master node.

Constraints

N/A

Range

  • true: The custom automation script runs only on the active Master nodes.
  • false: The custom automation script can run on all Master nodes.

Default Value

false

fail_action

Yes

String

Definition

Whether to continue executing subsequent scripts and creating a cluster after the custom automation script fails to be executed. You are advised to set this parameter to continue in the commissioning phase so the cluster can continue to be installed and started no matter whether the custom automation script is executed successfully.

Constraints

The scale-in operation cannot be undone. fail_action must be set to continue for the scripts that are executed after scale-in.

Range

  • continue: Continue to execute subsequent scripts.
  • errorout: Stop the action.

Default Value

N/A

action_stage

Yes

String

Definition

Time when a script is executed.

Constraints

N/A

Range

  • before_scale_out: before scale-out
  • before_scale_in: before scale-in
  • after_scale_out: after scale-out
  • after_scale_in: after scale-in

Default Value

N/A

Table 13 BootstrapScript

Parameter

Mandatory

Type

Description

name

Yes

String

Definition

Name of a bootstrap action script.

Constraints

N/A

Range

The names of bootstrap action scripts in the same cluster must be unique. The value can contain only numbers, letters, spaces, hyphens (-), and underscores (_) and cannot start with a space. The value can contain 1 to 64 characters.

Default Value

N/A

uri

Yes

String

Definition

Path of a bootstrap action script. Set this parameter to an OBS bucket path or a local VM path. OBS bucket path: Enter a script path, for example, enter the path of the public sample script provided by MRS. Example: obs://bootstrap/presto/presto-install.sh. If dualroles is installed, the parameter of the presto-install.sh script is dualroles. If worker is installed, the parameter of the presto-install.sh script is worker. Based on the Presto usage habit, you are advised to install dualroles on the active master nodes and worker on the core nodes. Local VM path: Enter a script path. The script path must start with a slash (/) and end with .sh.

Constraints

N/A

Range

N/A

Default Value

N/A

parameters

No

String

Definition

Bootstrap action script parameters

Constraints

N/A

Range

N/A

Default Value

N/A

nodes

Yes

Array of strings

Definition

Name of the node group where the bootstrap action script is executed

Constraints

N/A

Range

N/A

Default Value

N/A

active_master

No

Boolean

Definition

Whether the bootstrap action script runs only on active master nodes.

Constraints

N/A

Range

  • true: The bootstrap action script runs only on active Master nodes.
  • false: The bootstrap action script can run on all Master nodes.

Default Value

N/A

fail_action

Yes

String

Definition

Whether to continue executing subsequent scripts and creating a cluster after the bootstrap action script fails to execute. The default value is errorout, indicating that the action is stopped. Note: You are advised to set this parameter to continue in the commissioning phase so that the cluster can continue to be installed and started no matter whether the bootstrap action is successful.

Constraints

N/A

Range

  • continue: Continue to execute subsequent scripts.
  • errorout: Stop the action.

Default Value

errorout

before_component_start

No

Boolean

Definition

Time when the bootstrap action script is executed. Currently, the following two options are available: Before component start and After component start.

Constraints

N/A

Range

  • true: The bootstrap action script is executed before the component starts.
  • false: The bootstrap action script is executed after the component starts.

Default Value

false

start_time

No

Long

Definition

Execution time for a single bootstrap action script. The value is a Unix timestamp in seconds.

Constraints

N/A

Range

N/A

Default Value

N/A

state

No

String

Definition

Running state of an individual bootstrap action script.

Constraints

N/A

Range

  • PENDING: The script is suspended.
  • IN_PROGRESS: The script is being processed.
  • SUCCESS
  • FAILURE: The script fails to be executed.

Default Value

N/A

action_stages

No

Array of strings

Definition

Select the time when the bootstrap action script is executed.

Constraints

N/A

Range

  • BEFORE_COMPONENT_FIRST_START: before initial component starts
  • AFTER_COMPONENT_FIRST_START: after initial component starts
  • BEFORE_SCALE_IN: before scale-in
  • AFTER_SCALE_IN: after scale-in
  • BEFORE_SCALE_OUT: before scale-out
  • AFTER_SCALE_OUT: after scale-out

Default Value

N/A

Table 14 ComponentConfig

Parameter

Mandatory

Type

Description

component_name

Yes

String

Definition

Component name

Constraints

N/A

Range

N/A

Default Value

N/A

configs

No

Array of Config objects

Definition

Component configuration item list. For details about this parameter, see Table 15.

Constraints

The number of records cannot exceed 100.

Range

N/A

Default Value

N/A

Table 15 Config

Parameter

Mandatory

Type

Description

key

Yes

String

Definition

Configuration name. Only the configuration names displayed on the MRS component configuration page are supported.

Constraints

N/A

Range

N/A

Default Value

N/A

value

Yes

String

Definition

Configuration value.

Constraints

N/A

Range

N/A

Default Value

N/A

config_file_name

Yes

String

Definition

Configuration file name. Only the file names displayed on the MRS component configuration page are supported.

Constraints

N/A

Range

N/A

Default Value

N/A

Table 16 StepConfig

Parameter

Mandatory

Type

Description

job_execution

Yes

JobExecution object

Definition

Job parameter. For details about this parameter, see Table 17.

Constraints

N/A

Range

N/A

Default Value

N/A

Table 17 JobExecution

Parameter

Mandatory

Type

Description

job_type

Yes

String

Definition

Job type.

Constraints

N/A

Range

  • MapReduce: provides a distributed data processing model and execution environment capable of rapidly handling large-scale data in parallel. With MRS, you can submit MapReduce JAR programs.

  • SparkSubmit: allows you to submit Spark JAR and Spark Python programs and run Spark applications to compute and process user data.

  • SparkPython: SparkPython jobs are converted to SparkSubmit jobs for submission. On the MRS console, the job type is displayed as SparkSubmit. When calling an API to query the job list, select SparkSubmit.

  • HiveScript: an open-source data warehouse that runs on Hadoop. With MRS, you can submit HiveScript scripts for execution.

  • HiveSql: an open-source data warehouse that runs on Hadoop. With MRS, you can directly execute Hive SQL statements.

  • DistCp: a Hadoop tool used to efficiently import and export data between distributed file systems (such as HDFS).

  • SparkScript: allows you to submit SparkScript scripts and batch execute Spark SQL statements.

  • SparkSql: allows you to run SQL-like statements provided by Spark to query and analyze user data in real time.

  • Flink: a distributed big data processing engine that can perform stateful computations over both unbounded and bounded data streams.

Default Value

N/A

job_name

Yes

String

Definition

Job name.

Constraints

N/A

Range

A cluster name can contain only 1 to 64 characters. Only letters, digits, hyphens (-), and underscores (_) are allowed. Identical job names are allowed but not recommended.

Default Value

N/A

arguments

No

Array of strings

Definition

Key parameter for program execution. The parameter is specified by the function of the user's program. MRS is only responsible for loading the parameter.

Constraints

The value can contain a maximum of 150,000 characters. Special characters (;|&>'<$!"\) are not allowed. This parameter can be left blank.

Note:

  • When entering parameters with sensitive information, such as a login password, be aware that these values may appear in job details or logs. Exercise caution when using such inputs.
  • If you need to access files stored in OBS via the path starting with obs:// when submitting a HiveScript or HiveSQL job, search for the core.site.customized.configs parameter on the Hive service configuration page, add the endpoint configuration item (fs.obs.endpoint) of OBS, and set the value to the endpoint corresponding to OBS. For details, see Endpoints.

Range

N/A

Default Value

N/A

properties

No

Map<String,String>

Definition

Program system parameters, in the format of key/value pairs. Key indicates the parameter name, and Value indicates the parameter value.

Constraints

The value can contain a maximum of 2,048 characters. Special characters (;|&>'<$!\\) are not allowed. This parameter can be left blank.

Range

N/A

Default Value

N/A

Response Parameters

Status code: 200

Table 18 Response body parameter

Parameter

Type

Description

cluster_id

String

Definition

Cluster ID, which is returned by the system after the cluster is created.

Range

N/A

Example Request

Create an MRS 3.2.0-LTS.1 cluster where the custom management nodes and control nodes are the same nodes and submit a HiveScript job.

POST /v2/{project_id}/run-job-flow

{
  "cluster_version" : "MRS 3.1.0",
  "cluster_name" : "mrs_heshe_dm",
  "cluster_type" : "CUSTOM",
  "charge_info" : {
    "charge_mode" : "postPaid"
  },
  "region" : "",
  "availability_zone" : "",
  "vpc_name" : "vpc-37cd",
  "subnet_id" : "1f8c5ca6-1f66-4096-bb00-baf175954f6e",
  "subnet_name" : "subnet",
  "components" : "Hadoop,Spark2x,HBase,Hive,Hue,Loader,Kafka,Storm,Flume,Flink,Oozie,Ranger,Tez",
  "safe_mode" : "KERBEROS",
  "manager_admin_password" : "your password",
  "login_mode" : "PASSWORD",
  "node_root_password" : "your password",
  "mrs_ecs_default_agency" : "MRS_ECS_DEFAULT_AGENCY",
  "template_id" : "mgmt_control_combined_v2",
  "log_collection" : 1,
  "tags" : [ {
    "key" : "tag1",
    "value" : "111"
  }, {
    "key" : "tag2",
    "value" : "222"
  } ],
  "node_groups" : [ {
    "group_name" : "master_node_default_group",
    "node_num" : 3,
    "node_size" : "Sit3.4xlarge.4.linux.bigdata",
    "root_volume" : {
      "type" : "SAS",
      "size" : 480
    },
    "data_volume" : {
      "type" : "SAS",
      "size" : 600
    },
    "data_volume_count" : 1,
    "assigned_roles" : [ "OMSServer:1,2", "SlapdServer:1,2", "KerberosServer:1,2", "KerberosAdmin:1,2", "quorumpeer:1,2,3", "NameNode:2,3", "Zkfc:2,3", "JournalNode:1,2,3", "ResourceManager:2,3", "JobHistoryServer:2,3", "DBServer:1,3", "Hue:1,3", "LoaderServer:1,3", "MetaStore:1,2,3", "WebHCat:1,2,3", "HiveServer:1,2,3", "HMaster:2,3", "MonitorServer:1,2", "Nimbus:1,2", "UI:1,2", "JDBCServer2x:1,2,3", "JobHistory2x:2,3", "SparkResource2x:1,2,3", "oozie:2,3", "LoadBalancer:2,3", "TezUI:1,3", "TimelineServer:3", "RangerAdmin:1,2", "UserSync:2", "TagSync:2", "KerberosClient", "SlapdClient", "meta", "HSConsole:2,3", "FlinkResource:1,2,3", "DataNode:1,2,3", "NodeManager:1,2,3", "IndexServer2x:1,2", "ThriftServer:1,2,3", "RegionServer:1,2,3", "ThriftServer1:1,2,3", "RESTServer:1,2,3", "Broker:1,2,3", "Supervisor:1,2,3", "Logviewer:1,2,3", "Flume:1,2,3", "HSBroker:1,2,3" ]
  }, {
    "group_name" : "node_group_1",
    "node_num" : 3,
    "node_size" : "Sit3.4xlarge.4.linux.bigdata",
    "root_volume" : {
      "type" : "SAS",
      "size" : 480
    },
    "data_volume" : {
      "type" : "SAS",
      "size" : 600
    },
    "data_volume_count" : 1,
    "assigned_roles" : [ "DataNode", "NodeManager", "RegionServer", "Flume:1", "Broker", "Supervisor", "Logviewer", "HBaseIndexer", "KerberosClient", "SlapdClient", "meta", "HSBroker:1,2", "ThriftServer", "ThriftServer1", "RESTServer", "FlinkResource" ]
  }, {
    "group_name" : "node_group_2",
    "node_num" : 1,
    "node_size" : "Sit3.4xlarge.4.linux.bigdata",
    "root_volume" : {
      "type" : "SAS",
      "size" : 480
    },
    "data_volume" : {
      "type" : "SAS",
      "size" : 600
    },
    "data_volume_count" : 1,
    "assigned_roles" : [ "NodeManager", "KerberosClient", "SlapdClient", "meta", "FlinkResource" ]
  } ],
  "log_uri" : "obs://bucketTest/logs",
  "delete_when_no_steps" : true,
  "steps" : [ {
    "job_execution" : {
      "job_name" : "import_file",
      "job_type" : "DistCp",
      "arguments" : [ "obs://test/test.sql", "/user/hive/input" ]
    }
  }, {
    "job_execution" : {
      "job_name" : "hive_test",
      "job_type" : "HiveScript",
      "arguments" : [ "obs://test/hive/sql/HiveScript.sql" ]
    }
  } ]
}

Example Response

Status code: 200

Example successful response

{
  "cluster_id" : "da1592c2-bb7e-468d-9ac9-83246e95447a"
}

SDK Sample Code

The SDK sample code is as follows.

Create a custom cluster with combined management and control nodes, with cluster version MRS 3.1.0, and submit a HiveScript job.

  1
  2
  3
  4
  5
  6
  7
  8
  9
 10
 11
 12
 13
 14
 15
 16
 17
 18
 19
 20
 21
 22
 23
 24
 25
 26
 27
 28
 29
 30
 31
 32
 33
 34
 35
 36
 37
 38
 39
 40
 41
 42
 43
 44
 45
 46
 47
 48
 49
 50
 51
 52
 53
 54
 55
 56
 57
 58
 59
 60
 61
 62
 63
 64
 65
 66
 67
 68
 69
 70
 71
 72
 73
 74
 75
 76
 77
 78
 79
 80
 81
 82
 83
 84
 85
 86
 87
 88
 89
 90
 91
 92
 93
 94
 95
 96
 97
 98
 99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
package com.huaweicloud.sdk.test;

import com.huaweicloud.sdk.core.auth.ICredential;
import com.huaweicloud.sdk.core.auth.BasicCredentials;
import com.huaweicloud.sdk.core.exception.ConnectionException;
import com.huaweicloud.sdk.core.exception.RequestTimeoutException;
import com.huaweicloud.sdk.core.exception.ServiceResponseException;
import com.huaweicloud.sdk.mrs.v2.region.MrsRegion;
import com.huaweicloud.sdk.mrs.v2.*;
import com.huaweicloud.sdk.mrs.v2.model.*;

import java.util.List;
import java.util.ArrayList;

public class RunJobFlowSolution {

    public static void main(String[] args) {
        // The AK and SK used for authentication are hard-coded or stored in plaintext, which has great security risks. It is recommended that the AK and SK be stored in ciphertext in configuration files or environment variables and decrypted during use to ensure security.
        // In this example, AK and SK are stored in environment variables for authentication. Before running this example, set environment variables CLOUD_SDK_AK and CLOUD_SDK_SK in the local environment
        String ak = System.getenv("CLOUD_SDK_AK");
        String sk = System.getenv("CLOUD_SDK_SK");
        String projectId = "{project_id}";

        ICredential auth = new BasicCredentials()
                .withProjectId(projectId)
                .withAk(ak)
                .withSk(sk);

        MrsClient client = MrsClient.newBuilder()
                .withCredential(auth)
                .withRegion(MrsRegion.valueOf("<YOUR REGION>"))
                .build();
        RunJobFlowRequest request = new RunJobFlowRequest();
        RunJobFlowCommand body = new RunJobFlowCommand();
        List<String> listJobExecutionArguments = new ArrayList<>();
        listJobExecutionArguments.add("obs://test/hive/sql/HiveScript.sql");
        JobExecution jobExecutionSteps = new JobExecution();
        jobExecutionSteps.withJobType("HiveScript")
            .withJobName("hive_test")
            .withArguments(listJobExecutionArguments);
        List<String> listJobExecutionArguments1 = new ArrayList<>();
        listJobExecutionArguments1.add("obs://test/test.sql");
        listJobExecutionArguments1.add("/user/hive/input");
        JobExecution jobExecutionSteps1 = new JobExecution();
        jobExecutionSteps1.withJobType("DistCp")
            .withJobName("import_file")
            .withArguments(listJobExecutionArguments1);
        List<StepConfig> listbodySteps = new ArrayList<>();
        listbodySteps.add(
            new StepConfig()
                .withJobExecution(jobExecutionSteps1)
        );
        listbodySteps.add(
            new StepConfig()
                .withJobExecution(jobExecutionSteps)
        );
        List<String> listNodeGroupsAssignedRoles = new ArrayList<>();
        listNodeGroupsAssignedRoles.add("NodeManager");
        listNodeGroupsAssignedRoles.add("KerberosClient");
        listNodeGroupsAssignedRoles.add("SlapdClient");
        listNodeGroupsAssignedRoles.add("meta");
        listNodeGroupsAssignedRoles.add("FlinkResource");
        Volume dataVolumeNodeGroups = new Volume();
        dataVolumeNodeGroups.withType("SAS")
            .withSize(600);
        Volume rootVolumeNodeGroups = new Volume();
        rootVolumeNodeGroups.withType("SAS")
            .withSize(480);
        List<String> listNodeGroupsAssignedRoles1 = new ArrayList<>();
        listNodeGroupsAssignedRoles1.add("DataNode");
        listNodeGroupsAssignedRoles1.add("NodeManager");
        listNodeGroupsAssignedRoles1.add("RegionServer");
        listNodeGroupsAssignedRoles1.add("Flume:1");
        listNodeGroupsAssignedRoles1.add("Broker");
        listNodeGroupsAssignedRoles1.add("Supervisor");
        listNodeGroupsAssignedRoles1.add("Logviewer");
        listNodeGroupsAssignedRoles1.add("HBaseIndexer");
        listNodeGroupsAssignedRoles1.add("KerberosClient");
        listNodeGroupsAssignedRoles1.add("SlapdClient");
        listNodeGroupsAssignedRoles1.add("meta");
        listNodeGroupsAssignedRoles1.add("HSBroker:1,2");
        listNodeGroupsAssignedRoles1.add("ThriftServer");
        listNodeGroupsAssignedRoles1.add("ThriftServer1");
        listNodeGroupsAssignedRoles1.add("RESTServer");
        listNodeGroupsAssignedRoles1.add("FlinkResource");
        Volume dataVolumeNodeGroups1 = new Volume();
        dataVolumeNodeGroups1.withType("SAS")
            .withSize(600);
        Volume rootVolumeNodeGroups1 = new Volume();
        rootVolumeNodeGroups1.withType("SAS")
            .withSize(480);
        List<String> listNodeGroupsAssignedRoles2 = new ArrayList<>();
        listNodeGroupsAssignedRoles2.add("OMSServer:1,2");
        listNodeGroupsAssignedRoles2.add("SlapdServer:1,2");
        listNodeGroupsAssignedRoles2.add("KerberosServer:1,2");
        listNodeGroupsAssignedRoles2.add("KerberosAdmin:1,2");
        listNodeGroupsAssignedRoles2.add("quorumpeer:1,2,3");
        listNodeGroupsAssignedRoles2.add("NameNode:2,3");
        listNodeGroupsAssignedRoles2.add("Zkfc:2,3");
        listNodeGroupsAssignedRoles2.add("JournalNode:1,2,3");
        listNodeGroupsAssignedRoles2.add("ResourceManager:2,3");
        listNodeGroupsAssignedRoles2.add("JobHistoryServer:2,3");
        listNodeGroupsAssignedRoles2.add("DBServer:1,3");
        listNodeGroupsAssignedRoles2.add("Hue:1,3");
        listNodeGroupsAssignedRoles2.add("LoaderServer:1,3");
        listNodeGroupsAssignedRoles2.add("MetaStore:1,2,3");
        listNodeGroupsAssignedRoles2.add("WebHCat:1,2,3");
        listNodeGroupsAssignedRoles2.add("HiveServer:1,2,3");
        listNodeGroupsAssignedRoles2.add("HMaster:2,3");
        listNodeGroupsAssignedRoles2.add("MonitorServer:1,2");
        listNodeGroupsAssignedRoles2.add("Nimbus:1,2");
        listNodeGroupsAssignedRoles2.add("UI:1,2");
        listNodeGroupsAssignedRoles2.add("JDBCServer2x:1,2,3");
        listNodeGroupsAssignedRoles2.add("JobHistory2x:2,3");
        listNodeGroupsAssignedRoles2.add("SparkResource2x:1,2,3");
        listNodeGroupsAssignedRoles2.add("oozie:2,3");
        listNodeGroupsAssignedRoles2.add("LoadBalancer:2,3");
        listNodeGroupsAssignedRoles2.add("TezUI:1,3");
        listNodeGroupsAssignedRoles2.add("TimelineServer:3");
        listNodeGroupsAssignedRoles2.add("RangerAdmin:1,2");
        listNodeGroupsAssignedRoles2.add("UserSync:2");
        listNodeGroupsAssignedRoles2.add("TagSync:2");
        listNodeGroupsAssignedRoles2.add("KerberosClient");
        listNodeGroupsAssignedRoles2.add("SlapdClient");
        listNodeGroupsAssignedRoles2.add("meta");
        listNodeGroupsAssignedRoles2.add("HSConsole:2,3");
        listNodeGroupsAssignedRoles2.add("FlinkResource:1,2,3");
        listNodeGroupsAssignedRoles2.add("DataNode:1,2,3");
        listNodeGroupsAssignedRoles2.add("NodeManager:1,2,3");
        listNodeGroupsAssignedRoles2.add("IndexServer2x:1,2");
        listNodeGroupsAssignedRoles2.add("ThriftServer:1,2,3");
        listNodeGroupsAssignedRoles2.add("RegionServer:1,2,3");
        listNodeGroupsAssignedRoles2.add("ThriftServer1:1,2,3");
        listNodeGroupsAssignedRoles2.add("RESTServer:1,2,3");
        listNodeGroupsAssignedRoles2.add("Broker:1,2,3");
        listNodeGroupsAssignedRoles2.add("Supervisor:1,2,3");
        listNodeGroupsAssignedRoles2.add("Logviewer:1,2,3");
        listNodeGroupsAssignedRoles2.add("Flume:1,2,3");
        listNodeGroupsAssignedRoles2.add("HSBroker:1,2,3");
        Volume dataVolumeNodeGroups2 = new Volume();
        dataVolumeNodeGroups2.withType("SAS")
            .withSize(600);
        Volume rootVolumeNodeGroups2 = new Volume();
        rootVolumeNodeGroups2.withType("SAS")
            .withSize(480);
        List<NodeGroupV2> listbodyNodeGroups = new ArrayList<>();
        listbodyNodeGroups.add(
            new NodeGroupV2()
                .withGroupName("master_node_default_group")
                .withNodeNum(3)
                .withNodeSize("Sit3.4xlarge.4.linux.bigdata")
                .withRootVolume(rootVolumeNodeGroups2)
                .withDataVolume(dataVolumeNodeGroups2)
                .withDataVolumeCount(1)
                .withAssignedRoles(listNodeGroupsAssignedRoles2)
        );
        listbodyNodeGroups.add(
            new NodeGroupV2()
                .withGroupName("node_group_1")
                .withNodeNum(3)
                .withNodeSize("Sit3.4xlarge.4.linux.bigdata")
                .withRootVolume(rootVolumeNodeGroups1)
                .withDataVolume(dataVolumeNodeGroups1)
                .withDataVolumeCount(1)
                .withAssignedRoles(listNodeGroupsAssignedRoles1)
        );
        listbodyNodeGroups.add(
            new NodeGroupV2()
                .withGroupName("node_group_2")
                .withNodeNum(1)
                .withNodeSize("Sit3.4xlarge.4.linux.bigdata")
                .withRootVolume(rootVolumeNodeGroups)
                .withDataVolume(dataVolumeNodeGroups)
                .withDataVolumeCount(1)
                .withAssignedRoles(listNodeGroupsAssignedRoles)
        );
        List<Tag> listbodyTags = new ArrayList<>();
        listbodyTags.add(
            new Tag()
                .withKey("tag1")
                .withValue("111")
        );
        listbodyTags.add(
            new Tag()
                .withKey("tag2")
                .withValue("222")
        );
        ChargeInfo chargeInfobody = new ChargeInfo();
        chargeInfobody.withChargeMode("postPaid");
        body.withSteps(listbodySteps);
        body.withDeleteWhenNoSteps(true);
        body.withLogUri("obs://bucketTest/logs");
        body.withNodeGroups(listbodyNodeGroups);
        body.withLogCollection(RunJobFlowCommand.LogCollectionEnum.NUMBER_1);
        body.withTags(listbodyTags);
        body.withTemplateId("mgmt_control_combined_v2");
        body.withMrsEcsDefaultAgency("MRS_ECS_DEFAULT_AGENCY");
        body.withNodeRootPassword("your password");
        body.withLoginMode("PASSWORD");
        body.withManagerAdminPassword("your password");
        body.withSafeMode("KERBEROS");
        body.withAvailabilityZone("");
        body.withComponents("Hadoop,Spark2x,HBase,Hive,Hue,Loader,Kafka,Storm,Flume,Flink,Oozie,Ranger,Tez");
        body.withSubnetName("subnet");
        body.withSubnetId("1f8c5ca6-1f66-4096-bb00-baf175954f6e");
        body.withVpcName("vpc-37cd");
        body.withRegion("");
        body.withChargeInfo(chargeInfobody);
        body.withClusterType("CUSTOM");
        body.withClusterName("mrs_heshe_dm");
        body.withClusterVersion("MRS 3.1.0");
        request.withBody(body);
        try {
            RunJobFlowResponse response = client.runJobFlow(request);
            System.out.println(response.toString());
        } catch (ConnectionException e) {
            e.printStackTrace();
        } catch (RequestTimeoutException e) {
            e.printStackTrace();
        } catch (ServiceResponseException e) {
            e.printStackTrace();
            System.out.println(e.getHttpStatusCode());
            System.out.println(e.getRequestId());
            System.out.println(e.getErrorCode());
            System.out.println(e.getErrorMsg());
        }
    }
}

Create a custom cluster with combined management and control nodes, with cluster version MRS 3.1.0, and submit a HiveScript job.

  1
  2
  3
  4
  5
  6
  7
  8
  9
 10
 11
 12
 13
 14
 15
 16
 17
 18
 19
 20
 21
 22
 23
 24
 25
 26
 27
 28
 29
 30
 31
 32
 33
 34
 35
 36
 37
 38
 39
 40
 41
 42
 43
 44
 45
 46
 47
 48
 49
 50
 51
 52
 53
 54
 55
 56
 57
 58
 59
 60
 61
 62
 63
 64
 65
 66
 67
 68
 69
 70
 71
 72
 73
 74
 75
 76
 77
 78
 79
 80
 81
 82
 83
 84
 85
 86
 87
 88
 89
 90
 91
 92
 93
 94
 95
 96
 97
 98
 99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
# coding: utf-8

import os
from huaweicloudsdkcore.auth.credentials import BasicCredentials
from huaweicloudsdkmrs.v2.region.mrs_region import MrsRegion
from huaweicloudsdkcore.exceptions import exceptions
from huaweicloudsdkmrs.v2 import *

if __name__ == "__main__":
    # The AK and SK used for authentication are hard-coded or stored in plaintext, which has great security risks. It is recommended that the AK and SK be stored in ciphertext in configuration files or environment variables and decrypted during use to ensure security.
    # In this example, AK and SK are stored in environment variables for authentication. Before running this example, set environment variables CLOUD_SDK_AK and CLOUD_SDK_SK in the local environment
    ak = os.environ["CLOUD_SDK_AK"]
    sk = os.environ["CLOUD_SDK_SK"]
    projectId = "{project_id}"

    credentials = BasicCredentials(ak, sk, projectId)

    client = MrsClient.new_builder() \
        .with_credentials(credentials) \
        .with_region(MrsRegion.value_of("<YOUR REGION>")) \
        .build()

    try:
        request = RunJobFlowRequest()
        listArgumentsJobExecution = [
            "obs://test/hive/sql/HiveScript.sql"
        ]
        jobExecutionSteps = JobExecution(
            job_type="HiveScript",
            job_name="hive_test",
            arguments=listArgumentsJobExecution
        )
        listArgumentsJobExecution1 = [
            "obs://test/test.sql",
            "/user/hive/input"
        ]
        jobExecutionSteps1 = JobExecution(
            job_type="DistCp",
            job_name="import_file",
            arguments=listArgumentsJobExecution1
        )
        listStepsbody = [
            StepConfig(
                job_execution=jobExecutionSteps1
            ),
            StepConfig(
                job_execution=jobExecutionSteps
            )
        ]
        listAssignedRolesNodeGroups = [
            "NodeManager",
            "KerberosClient",
            "SlapdClient",
            "meta",
            "FlinkResource"
        ]
        dataVolumeNodeGroups = Volume(
            type="SAS",
            size=600
        )
        rootVolumeNodeGroups = Volume(
            type="SAS",
            size=480
        )
        listAssignedRolesNodeGroups1 = [
            "DataNode",
            "NodeManager",
            "RegionServer",
            "Flume:1",
            "Broker",
            "Supervisor",
            "Logviewer",
            "HBaseIndexer",
            "KerberosClient",
            "SlapdClient",
            "meta",
            "HSBroker:1,2",
            "ThriftServer",
            "ThriftServer1",
            "RESTServer",
            "FlinkResource"
        ]
        dataVolumeNodeGroups1 = Volume(
            type="SAS",
            size=600
        )
        rootVolumeNodeGroups1 = Volume(
            type="SAS",
            size=480
        )
        listAssignedRolesNodeGroups2 = [
            "OMSServer:1,2",
            "SlapdServer:1,2",
            "KerberosServer:1,2",
            "KerberosAdmin:1,2",
            "quorumpeer:1,2,3",
            "NameNode:2,3",
            "Zkfc:2,3",
            "JournalNode:1,2,3",
            "ResourceManager:2,3",
            "JobHistoryServer:2,3",
            "DBServer:1,3",
            "Hue:1,3",
            "LoaderServer:1,3",
            "MetaStore:1,2,3",
            "WebHCat:1,2,3",
            "HiveServer:1,2,3",
            "HMaster:2,3",
            "MonitorServer:1,2",
            "Nimbus:1,2",
            "UI:1,2",
            "JDBCServer2x:1,2,3",
            "JobHistory2x:2,3",
            "SparkResource2x:1,2,3",
            "oozie:2,3",
            "LoadBalancer:2,3",
            "TezUI:1,3",
            "TimelineServer:3",
            "RangerAdmin:1,2",
            "UserSync:2",
            "TagSync:2",
            "KerberosClient",
            "SlapdClient",
            "meta",
            "HSConsole:2,3",
            "FlinkResource:1,2,3",
            "DataNode:1,2,3",
            "NodeManager:1,2,3",
            "IndexServer2x:1,2",
            "ThriftServer:1,2,3",
            "RegionServer:1,2,3",
            "ThriftServer1:1,2,3",
            "RESTServer:1,2,3",
            "Broker:1,2,3",
            "Supervisor:1,2,3",
            "Logviewer:1,2,3",
            "Flume:1,2,3",
            "HSBroker:1,2,3"
        ]
        dataVolumeNodeGroups2 = Volume(
            type="SAS",
            size=600
        )
        rootVolumeNodeGroups2 = Volume(
            type="SAS",
            size=480
        )
        listNodeGroupsbody = [
            NodeGroupV2(
                group_name="master_node_default_group",
                node_num=3,
                node_size="Sit3.4xlarge.4.linux.bigdata",
                root_volume=rootVolumeNodeGroups2,
                data_volume=dataVolumeNodeGroups2,
                data_volume_count=1,
                assigned_roles=listAssignedRolesNodeGroups2
            ),
            NodeGroupV2(
                group_name="node_group_1",
                node_num=3,
                node_size="Sit3.4xlarge.4.linux.bigdata",
                root_volume=rootVolumeNodeGroups1,
                data_volume=dataVolumeNodeGroups1,
                data_volume_count=1,
                assigned_roles=listAssignedRolesNodeGroups1
            ),
            NodeGroupV2(
                group_name="node_group_2",
                node_num=1,
                node_size="Sit3.4xlarge.4.linux.bigdata",
                root_volume=rootVolumeNodeGroups,
                data_volume=dataVolumeNodeGroups,
                data_volume_count=1,
                assigned_roles=listAssignedRolesNodeGroups
            )
        ]
        listTagsbody = [
            Tag(
                key="tag1",
                value="111"
            ),
            Tag(
                key="tag2",
                value="222"
            )
        ]
        chargeInfobody = ChargeInfo(
            charge_mode="postPaid"
        )
        request.body = RunJobFlowCommand(
            steps=listStepsbody,
            delete_when_no_steps=True,
            log_uri="obs://bucketTest/logs",
            node_groups=listNodeGroupsbody,
            log_collection=1,
            tags=listTagsbody,
            template_id="mgmt_control_combined_v2",
            mrs_ecs_default_agency="MRS_ECS_DEFAULT_AGENCY",
            node_root_password="your password",
            login_mode="PASSWORD",
            manager_admin_password="your password",
            safe_mode="KERBEROS",
            availability_zone="",
            components="Hadoop,Spark2x,HBase,Hive,Hue,Loader,Kafka,Storm,Flume,Flink,Oozie,Ranger,Tez",
            subnet_name="subnet",
            subnet_id="1f8c5ca6-1f66-4096-bb00-baf175954f6e",
            vpc_name="vpc-37cd",
            region="",
            charge_info=chargeInfobody,
            cluster_type="CUSTOM",
            cluster_name="mrs_heshe_dm",
            cluster_version="MRS 3.1.0"
        )
        response = client.run_job_flow(request)
        print(response)
    except exceptions.ClientRequestException as e:
        print(e.status_code)
        print(e.request_id)
        print(e.error_code)
        print(e.error_msg)

Create a custom cluster with combined management and control nodes, with cluster version MRS 3.1.0, and submit a HiveScript job.

  1
  2
  3
  4
  5
  6
  7
  8
  9
 10
 11
 12
 13
 14
 15
 16
 17
 18
 19
 20
 21
 22
 23
 24
 25
 26
 27
 28
 29
 30
 31
 32
 33
 34
 35
 36
 37
 38
 39
 40
 41
 42
 43
 44
 45
 46
 47
 48
 49
 50
 51
 52
 53
 54
 55
 56
 57
 58
 59
 60
 61
 62
 63
 64
 65
 66
 67
 68
 69
 70
 71
 72
 73
 74
 75
 76
 77
 78
 79
 80
 81
 82
 83
 84
 85
 86
 87
 88
 89
 90
 91
 92
 93
 94
 95
 96
 97
 98
 99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
package main

import (
	"fmt"
	"github.com/huaweicloud/huaweicloud-sdk-go-v3/core/auth/basic"
    mrs "github.com/huaweicloud/huaweicloud-sdk-go-v3/services/mrs/v2"
	"github.com/huaweicloud/huaweicloud-sdk-go-v3/services/mrs/v2/model"
    region "github.com/huaweicloud/huaweicloud-sdk-go-v3/services/mrs/v2/region"
)

func main() {
    // The AK and SK used for authentication are hard-coded or stored in plaintext, which has great security risks. It is recommended that the AK and SK be stored in ciphertext in configuration files or environment variables and decrypted during use to ensure security.
    // In this example, AK and SK are stored in environment variables for authentication. Before running this example, set environment variables CLOUD_SDK_AK and CLOUD_SDK_SK in the local environment
    ak := os.Getenv("CLOUD_SDK_AK")
    sk := os.Getenv("CLOUD_SDK_SK")
    projectId := "{project_id}"

    auth, err := basic.NewCredentialsBuilder().
        WithAk(ak).
        WithSk(sk).
        WithProjectId(projectId).
        SafeBuild()

    if err != nil {
        fmt.Println(err)
        return
    }

    hcClient, err := mrs.MrsClientBuilder().
         WithRegion(region.ValueOf("<YOUR REGION>")).
         WithCredential(auth).
         SafeBuild()


    if err != nil {
        fmt.Println(err)
        return
    }

    client := mrs.NewMrsClient(hcClient)

    request := &model.RunJobFlowRequest{}
	var listArgumentsJobExecution = []string{
        "obs://test/hive/sql/HiveScript.sql",
    }
	jobExecutionSteps := &model.JobExecution{
		JobType: "HiveScript",
		JobName: "hive_test",
		Arguments: &listArgumentsJobExecution,
	}
	var listArgumentsJobExecution1 = []string{
        "obs://test/test.sql",
	    "/user/hive/input",
    }
	jobExecutionSteps1 := &model.JobExecution{
		JobType: "DistCp",
		JobName: "import_file",
		Arguments: &listArgumentsJobExecution1,
	}
	var listStepsbody = []model.StepConfig{
        {
            JobExecution: jobExecutionSteps1,
        },
        {
            JobExecution: jobExecutionSteps,
        },
    }
	var listAssignedRolesNodeGroups = []string{
        "NodeManager",
	    "KerberosClient",
	    "SlapdClient",
	    "meta",
	    "FlinkResource",
    }
	dataVolumeNodeGroups := &model.Volume{
		Type: "SAS",
		Size: int32(600),
	}
	rootVolumeNodeGroups := &model.Volume{
		Type: "SAS",
		Size: int32(480),
	}
	var listAssignedRolesNodeGroups1 = []string{
        "DataNode",
	    "NodeManager",
	    "RegionServer",
	    "Flume:1",
	    "Broker",
	    "Supervisor",
	    "Logviewer",
	    "HBaseIndexer",
	    "KerberosClient",
	    "SlapdClient",
	    "meta",
	    "HSBroker:1,2",
	    "ThriftServer",
	    "ThriftServer1",
	    "RESTServer",
	    "FlinkResource",
    }
	dataVolumeNodeGroups1 := &model.Volume{
		Type: "SAS",
		Size: int32(600),
	}
	rootVolumeNodeGroups1 := &model.Volume{
		Type: "SAS",
		Size: int32(480),
	}
	var listAssignedRolesNodeGroups2 = []string{
        "OMSServer:1,2",
	    "SlapdServer:1,2",
	    "KerberosServer:1,2",
	    "KerberosAdmin:1,2",
	    "quorumpeer:1,2,3",
	    "NameNode:2,3",
	    "Zkfc:2,3",
	    "JournalNode:1,2,3",
	    "ResourceManager:2,3",
	    "JobHistoryServer:2,3",
	    "DBServer:1,3",
	    "Hue:1,3",
	    "LoaderServer:1,3",
	    "MetaStore:1,2,3",
	    "WebHCat:1,2,3",
	    "HiveServer:1,2,3",
	    "HMaster:2,3",
	    "MonitorServer:1,2",
	    "Nimbus:1,2",
	    "UI:1,2",
	    "JDBCServer2x:1,2,3",
	    "JobHistory2x:2,3",
	    "SparkResource2x:1,2,3",
	    "oozie:2,3",
	    "LoadBalancer:2,3",
	    "TezUI:1,3",
	    "TimelineServer:3",
	    "RangerAdmin:1,2",
	    "UserSync:2",
	    "TagSync:2",
	    "KerberosClient",
	    "SlapdClient",
	    "meta",
	    "HSConsole:2,3",
	    "FlinkResource:1,2,3",
	    "DataNode:1,2,3",
	    "NodeManager:1,2,3",
	    "IndexServer2x:1,2",
	    "ThriftServer:1,2,3",
	    "RegionServer:1,2,3",
	    "ThriftServer1:1,2,3",
	    "RESTServer:1,2,3",
	    "Broker:1,2,3",
	    "Supervisor:1,2,3",
	    "Logviewer:1,2,3",
	    "Flume:1,2,3",
	    "HSBroker:1,2,3",
    }
	dataVolumeNodeGroups2 := &model.Volume{
		Type: "SAS",
		Size: int32(600),
	}
	rootVolumeNodeGroups2 := &model.Volume{
		Type: "SAS",
		Size: int32(480),
	}
	dataVolumeCountNodeGroups:= int32(1)
	dataVolumeCountNodeGroups1:= int32(1)
	dataVolumeCountNodeGroups2:= int32(1)
	var listNodeGroupsbody = []model.NodeGroupV2{
        {
            GroupName: "master_node_default_group",
            NodeNum: int32(3),
            NodeSize: "Sit3.4xlarge.4.linux.bigdata",
            RootVolume: rootVolumeNodeGroups2,
            DataVolume: dataVolumeNodeGroups2,
            DataVolumeCount: &dataVolumeCountNodeGroups,
            AssignedRoles: &listAssignedRolesNodeGroups2,
        },
        {
            GroupName: "node_group_1",
            NodeNum: int32(3),
            NodeSize: "Sit3.4xlarge.4.linux.bigdata",
            RootVolume: rootVolumeNodeGroups1,
            DataVolume: dataVolumeNodeGroups1,
            DataVolumeCount: &dataVolumeCountNodeGroups1,
            AssignedRoles: &listAssignedRolesNodeGroups1,
        },
        {
            GroupName: "node_group_2",
            NodeNum: int32(1),
            NodeSize: "Sit3.4xlarge.4.linux.bigdata",
            RootVolume: rootVolumeNodeGroups,
            DataVolume: dataVolumeNodeGroups,
            DataVolumeCount: &dataVolumeCountNodeGroups2,
            AssignedRoles: &listAssignedRolesNodeGroups,
        },
    }
	var listTagsbody = []model.Tag{
        {
            Key: "tag1",
            Value: "111",
        },
        {
            Key: "tag2",
            Value: "222",
        },
    }
	chargeInfobody := &model.ChargeInfo{
		ChargeMode: "postPaid",
	}
	deleteWhenNoStepsRunJobFlowCommand:= true
	logUriRunJobFlowCommand:= "obs://bucketTest/logs"
	logCollectionRunJobFlowCommand:= model.GetRunJobFlowCommandLogCollectionEnum().E_1
	templateIdRunJobFlowCommand:= "mgmt_control_combined_v2"
	mrsEcsDefaultAgencyRunJobFlowCommand:= "MRS_ECS_DEFAULT_AGENCY"
	nodeRootPasswordRunJobFlowCommand:= "your password"
	subnetIdRunJobFlowCommand:= "1f8c5ca6-1f66-4096-bb00-baf175954f6e"
	request.Body = &model.RunJobFlowCommand{
		Steps: listStepsbody,
		DeleteWhenNoSteps: &deleteWhenNoStepsRunJobFlowCommand,
		LogUri: &logUriRunJobFlowCommand,
		NodeGroups: listNodeGroupsbody,
		LogCollection: &logCollectionRunJobFlowCommand,
		Tags: &listTagsbody,
		TemplateId: &templateIdRunJobFlowCommand,
		MrsEcsDefaultAgency: &mrsEcsDefaultAgencyRunJobFlowCommand,
		NodeRootPassword: &nodeRootPasswordRunJobFlowCommand,
		LoginMode: "PASSWORD",
		ManagerAdminPassword: "your password",
		SafeMode: "KERBEROS",
		AvailabilityZone: "",
		Components: "Hadoop,Spark2x,HBase,Hive,Hue,Loader,Kafka,Storm,Flume,Flink,Oozie,Ranger,Tez",
		SubnetName: "subnet",
		SubnetId: &subnetIdRunJobFlowCommand,
		VpcName: "vpc-37cd",
		Region: "",
		ChargeInfo: chargeInfobody,
		ClusterType: "CUSTOM",
		ClusterName: "mrs_heshe_dm",
		ClusterVersion: "MRS 3.1.0",
	}
	response, err := client.RunJobFlow(request)
	if err == nil {
        fmt.Printf("%+v\n", response)
    } else {
        fmt.Println(err)
    }
}

For SDK sample code of more programming languages, see the Sample Code tab in API Explorer. SDK sample code can be automatically generated.

Status Codes

Status Code

Description

200

Example response for a successful request.

Error Codes

See Error Codes.