Updated on 2026-02-25 GMT+08:00

Creating a Hudi Table

Notes

  • DataArts Studio can execute Spark SQLs through proxy connections or API connections. When using API connections to execute Hudi-related commands, add --conf spark.support.hudi=true to the job running parameters.
  • The field names in the Hudi table cannot be capitalized because the Spark, Flink, Hive, and HetuEngine engines have case compatibility differences.
  • A Hudi table must have a primary key. The value of the primary key field cannot contain commas (,), colons (:), null, or empty values. The value cannot be changed after the table is created..
  • The Hudi table must contain the precombine field. The precombine field cannot contain null or empty values. precombine can be set to only one column in the Hudi table and cannot be modified after the table is created. Note that this field involves the update logic of the Hudi table. The data is updated only when the value of the column in the new data is greater than or equal to that of the column in the old data.
  • The index of the Hudi table has a default value. In versions earlier than MRS 3.3.1, the BLOOM index is used. In MRS 3.5.0 and later versions, the SIMPLE index is used. You are advised to specify an index suitable for the service when creating a table. The index cannot be modified after the table is created.
  • If bucket indexes are used, estimate the number of buckets by referring to Determining the Table Index. The number of buckets cannot be changed after the table is created.
  • If the location is specified in the table creation statement, the table is a foreign table. If the location is not specified, the table is an internal table. If you run drop table on an internal table, the table and data storage directory on Hive are deleted. If you run drop table on a foreign table, only the table on Hive is deleted. To delete the table and data storage directory on Hive, run drop table table name purge. Exercise caution when perform this operation.

DataArts Studio/Spark SQL

  1. Hudi table with a single primary key.
    create table hudi_table (
    id int,
    name string,
    price double
    ) using hudi
    options (
    type  = 'cow',
    primaryKey = 'id', --The primary key must be specified.
    preCombineField = 'id', -- The precombine field must be specified. Generally, the precombine field is set to the same field as the primary key to implement update by primary key.
    hoodie.index.type = 'SIMPLE' -- If this parameter is not specified, the default index is used.
    );
  2. Multi-primary key Hudi table.
    create table hudi_table (
    id1 int,
    id2 int,
    name string,
    price double
    ) using hudi
    options (
    type  = 'mor',
    primaryKey = 'id1,id2', --The primary key must be specified. The number of composite primary keys is not limited and the primary keys are separated by commas (,).
    preCombineField = 'id1', -- The precombine field must be specified. Only one column can be set for the precombine field.
    hoodie.index.type = 'BLOOM' -- If this parameter is not specified, the default index is used.
    );
  3. Hudi table with the BUCKET index.
    create table hudi_table (
    id1 int,
    id2 int,
    name string,
    price double
    ) using hudi
    options (
    type  = 'mor',
    primaryKey = 'id1,id2', --The primary key must be specified. The number of composite primary keys is not limited and the primary keys are separated by commas (,).
    preCombineField = 'id1', -- The precombine field must be specified. Only one column can be set for the precombine field.
    hoodie.index.type = 'BUCKET', -- must be specified.
    hoodie.bucket.index.num.buckets = '5', -- Mandatory. The number of buckets must be estimated by referring to section 6.2.2.
    hoodie.bucket.index.hash.field = 'id1,id2' -- is optional. The hash field of the bucket index is the same as that of the primary key by default. Generally, you do not need to set it.
    );
  4. Partitioned table.
    create table hudi_table (
    id1 int,
    id2 int,
    par1 int,
    par2 int,
    name string,
    price double
    ) using hudi
    options (
    ......
    ) partitioned by (par1, par2);
  5. Foreign table.
    create table hudi_table (
    id1 int,
    id2 int,
    par1 int,
    par2 int,
    name string,
    price double
    ) using hudi
    options (
    ......
    ) partitioned by (par1, par2) location  "hdfs://.../hudi_table"; -- HDFS path or OBS path

CDM/CDL

CDM/CDL migration tools are usually used to map source tables to Hudi tables on the GUI. Pay attention to the following points when creating a CDM/CDL job:

  • You must comply with the rules in Notes regardless of the platform.
  • Note that the Hudi table fields cannot be in uppercase because the source table of the migration tool is usually from the traditional data warehouse and the fields in the source table are usually in uppercase. As a result, uppercase letters are used when the Hudi table is created.