Help Center/ DataArts Studio/ User Guide/ DataArts Migration (CDM Jobs)/ Creating a Job in a CDM Cluster/ Reading and Writing Dirty Data from Different Sources
Updated on 2026-08-13 GMT+08:00

Reading and Writing Dirty Data from Different Sources

During data integration, dirty data may be generated due to data type mismatch, incorrect formats, or constraints and conflicts. CDM can record and write dirty data to a specified path for analysis and processing.

The following table lists whether dirty data from different sources can be read and written.

Table 1 Reading and writing of dirty data from different sources

Data Source

Read

Write

MySQL

√

√

PostgreSQL

√

√

SQL Server

√

√

SAP HANA

√

√

Oracle

√

√

GBase

√

√

GaussDB

√

√

DWS

√

√

DLI

x

√

MRS Hive and Apache Hive

√

x

MRS Hudi

√

x

MRS ClickHouse and Apache ClickHouse

√

x

Doris

√

√

MRS HBase

√

√

MongoDB

√

√

Redis

√

√

Elasticsearch

√

√

DMS Kafka

x

x

LTS

x

x

Apache RocketMQ

x

x

RestClient

x

x

Apache HDFS

√

x

OBS

√

x

FTP

√

x

SFTP

√

x