Updated on 2026-02-25 GMT+08:00

Determining the Table Index

Table index description

Index Type

Index Features

Supported Engine

Preferred Scenario

SIMPLE

  • Partition-level update: When data with the same primary key is written to different partitions, the update is not triggered. As a result, duplicate data exists.
  • Execute a join to complete the update, which consumes more memory.

Spark

  • COW tables, batch scenario.

BUCKET

  • Each row of data is hashed and stored in the corresponding bucket based on the configured number of buckets, with the fastest write speed.
  • There is no limit on the data volume, leading to outstanding performance in big data scenarios. Bucketing can split data and effectively control the number of files.
  • Compatible with multiple engines. This index must be used when Flink and Spark operate the same Hudi table at the same time.

Spark/Flink

  • MOR table, stream scenario, real-time write
  • MOR tables, batch scenario
  • COW tables, append scenario without update, real time write. This index applies to scenarios that have requirements on write performance and point query. However, a large number of small files are generated during appending. Therefore, partition filtering and bucket filtering are needed. This use scenario only supported by certain MRS versions. Consult the service support team before performing the operation.
  • COW tables, always insert overwrite.

BLOOM

  • Partition-level update: When data with the same primary key is written to different partitions, the update is not triggered. As a result, duplicate data exists.
  • The performance of BLOOM indexes in large data volume scenarios is influenced by the amount of data to be updated. The larger the proportion of data updates, the poorer the performance. BLOOM indexes are not recommended in large data volume scenarios because a large number of files are generated.

Spark

  • COW tables, batch scenario, less than 20% data to be updated
  • MOR tables, batch scenario, less than 20% data to be updated

GLOBAL_BLOOM

  • Table-level update. Data with the same primary key is also updated when it is written to different partitions.
  • The performance is poor is not recommended in large data volume scenarios. Therefore, this index is not recommended.

Spark

  • Less than one million data volume, global deduplication scenario

GLOBAL_SIMPLE

  • Table-level update. Data with the same primary key is also updated when it is written to different partitions.
  • The performance is poor is not recommended in large data volume scenarios. Therefore, this index is not recommended.

Spark

  • Less than one million data volume, global deduplication scenario

Typical Scenarios

  • COW tables are always written with INSERT OVERWRITE, and BUCKET indexes can be used.
  • When COW tables are written with INSERT INTO, exercise caution when using BUCKET indexes because BUCKET indexes may cause incremental data and all BUCKET buckets need to update. As a result, the write speed becomes slower. If the data volume is tens of thousands or millions, you can select SIMPLE or BUCKET for COW tables. If the data volume is more than 10 million, SIMPLE is recommended.
  • When MOR tables are written with INSERT INTO, BUCKET indexes are recommended. BUCKET indexes are applicable to write and read across multiple engines and scenarios with a large amount of data. However, compaction needs to be performed periodically.
  • Writing MOR tables with INSERT OVERWRITE is not recommended. You can use INSERT OVERWRITE and BUCKET indexes for COW tables.

Estimating the Number of Buckets in the BUCKET Index Table

The number of buckets of a Hudi table must be determined during table creation and cannot be changed later. If the number of buckets is not properly set, serious performance problems may occur. You must perform the following steps to estimate the number of buckets:

  • Non-partitioned tables
    1. The estimated total number of data records in the Hudi table is A, which cannot equal to the number of existing data records. Consider the growth rate of the Hudi table in the next five years. For example, the total data records of the Hudi table will increase to A in the next five years.
    2. Check the size of a single data record in the Hudi table, which is B (in KB). Use limit 100 to randomly query 100 service data records in the source table and save the 100 service data records to a TXT file. B = TXT file size (in KB) / 100.
    3. Check the data volume C (in GB) of the Hudi table before compression. C = A x B / 1024 / 1024.
    4. Estimate the number of buckets of the Hudi table D. D = MAX (rounded up (C/2 x 1.5), 4).

    D = MAX (rounded up (C/2 x 1.5), 4). In this formula, 2 indicates that 2 GB data is stored in a bucket, and 1.5 indicates that 1.5 times buckets are reserved for the non-partitioned table.

  • Partitioned tables
    1. Estimate the total number of data records in a single partition of the Hudi table as A. That is, in the next five years , the total number of data records in a single partition will increase to A. For example, if data is partitioned by day, you need to consider the amount of data generated on special holidays in the next few years. If data is partitioned by year, you need to consider the service data growth in the next few years.
    2. Check the size of a single data record in the Hudi table, which is B (in KB). Use limit 100 to randomly query 100 service data records in the source table and save the 100 service data records to a TXT file. B = TXT file size (in KB) / 100.
    3. Check the data volume C (in GB) of the Hudi table before compression. C = A x B / 1024 / 1024.
    4. Estimate the number of buckets of the Hudi table D. D = MAX (rounded up (C/2), 1).
    1. The number of buckets must be estimated based on the uncompressed data volume instead of the size of the compressed file in the source table, for example, the Parquet file.
    2. It is best to set an even number of buckets, with a minimum of 4 for non-partitioned tables and a minimum of 1 for partitioned tables.