- What's New
- Function Overview
- Service Overview
- Data Governance Methodology
- Preparations
- Getting Started
-
User Guide
- DataArts Studio development process
-
Buying and Configuring a DataArts Studio Instance
- Buying a DataArts Studio Instance
-
Buying a DataArts Studio Incremental Package
- Introduction to Incremental Packages
- Buying a DataArts Migration Incremental Package
- Buying a DataArts Migration Resource Group Incremental Package
- Buying a DataArts DataService Exclusive Cluster Incremental Package
- Buying an Incremental Package for Job Node Scheduling Times/Day
- Buying an Incremental Package for Technical Asset Quantity
- Buying an Incremental Package for Data Model Quantity
- Accessing the DataArts Studio Instance Console
- Creating and Configuring a Workspace in Simple Mode
- (Optional) Creating and Using a Workspace in Enterprise Mode
- Managing DataArts Studio Resources
- Authorizing Users to Use DataArts Studio
-
Management Center
- Data Sources Supported by DataArts Studio
- Creating a DataArts Studio Data Connection
-
Configuring DataArts Studio Data Connection Parameters
- DWS Connection Parameters
- DLI Connection Parameters
- MRS Hive Connection Parameters
- MRS HBase Connection Parameters
- MRS Kafka Connection Parameters
- MRS Spark Connection Parameters
- MRS ClickHouse Connection Parameters
- MRS Hetu Connection Parameters
- MRS Impala Connection Parameters
- MRS Ranger Connection Parameters
- MRS Presto Connection Parameters
- Doris Connection Parameters
- OpenSource ClickHouse Connection Parameters
- RDS Connection Parameters
- Oracle Connection Parameters
- DIS Connection Parameters
- Host Connection Parameters
- Rest Client Connection Parameters
- Redis Connection Parameters
- SAP HANA Connection Parameters
- LTS Connection Parameters
- Configuring DataArts Studio Resource Migration
- Configuring Environment Isolation for a DataArts Studio Workspace in Enterprise Mode
- Typical Scenarios for Using Management Center
-
DataArts Migration (CDM Jobs)
- Overview
- Notes and Constraints
- Supported Data Sources
- Creating and Managing a CDM Cluster
-
Creating a Link in a CDM Cluster
- Creating a Link Between CDM and a Data Source
-
Configuring Link Parameters
- OBS Link Parameters
- PostgreSQL/SQLServer Link Parameters
- GaussDB(DWS) Link Parameters
- RDS for MySQL/MySQL Database Link Parameters
- Oracle Database Link Parameters
- DLI Link Parameters
- Hive Link Parameters
- HBase Link Parameters
- HDFS Link Parameters
- FTP/SFTP Link Parameters
- Redis Link Parameters
- DDS Link Parameters
- CloudTable Link Parameters
- MongoDB Link Parameters
- Cassandra Link Parameters
- DIS Link Parameters
- Kafka Link Parameters
- DMS Kafka Link Parameters
- CSS Link Parameters
- Elasticsearch Link Parameters
- Dameng Database Link Parameters
- SAP HANA Link Parameters
- Shard Link Parameters
- MRS Hudi Link Parameters
- MRS ClickHouse Link Parameters
- ShenTong Database Link Parameters
- CloudTable OpenTSDB Link Parameters
- GBASE Link Parameters
- YASHAN Link Parameters
- Uploading a CDM Link Driver
- Creating a Hadoop Cluster Configuration
-
Creating a Job in a CDM Cluster
- Table/File Migration Jobs
- Creating an Entire Database Migration Job
-
Configuring CDM Source Job Parameters
- From OBS
- From HDFS
- From HBase/CloudTable
- From Hive
- From DLI
- From FTP/SFTP
- From HTTP
- From PostgreSQL/SQL Server
- From DWS
- From SAP HANA
- From MySQL
- From Oracle
- From a Database Shard
- From MongoDB/DDS
- From Redis
- From DIS
- From Kafka/DMS Kafka
- From Elasticsearch or CSS
- From OpenTSDB
- From MRS Hudi
- From MRS ClickHouse
- From a ShenTong Database
- From a Dameng Database
- From YASHAN
- Configuring CDM Destination Job Parameters
- Configuring CDM Job Field Mapping
- Configuring a Scheduled CDM Job
- Managing CDM Job Configuration
- Managing a CDM Job
- Managing CDM Jobs
- Using Macro Variables of Date and Time
- Improving Migration Performance
-
Key Operation Guide
- Incremental Migration
- Migration in Transaction Mode
- Encryption and Decryption During File Migration
- MD5 Verification
- Configuring Field Converters
- Adding Fields
- Migrating Files with Specified Names
- Regular Expressions for Separating Semi-structured Text
- Recording the Time When Data Is Written to the Database
- File Formats
- Converting Unsupported Data Types
- Auto Table Creation
-
Tutorials
- Creating an MRS Hive Link
- Creating a MySQL Link
- Migrating Data from MySQL to MRS Hive
- Migrating Data from MySQL to OBS
- Migrating Data from MySQL to DWS
- Migrating an Entire MySQL Database to RDS
- Migrating Data from Oracle to CSS
- Migrating Data from Oracle to DWS
- Migrating Data from OBS to CSS
- Migrating Data from OBS to DLI
- Migrating Data from MRS HDFS to OBS
- Migrating the Entire Elasticsearch Database to CSS
- Error Codes
- DataArts Migration (Offline Jobs)
-
DataArts Migration (Real-Time Jobs)
- Overview of Real-Time Jobs
- Supported Data Sources
- Check Before Use
-
Enabling Network Communications
- Database Deployed in an On-premises IDC
- Database Deployed on Another Cloud
-
Database Deployed on Huawei Cloud
- Enabling Network Communications Directly for the Same Region and Tenant
- Using a VPC Peering Connection to Enable Network Communications for the Same Region but Different Tenants
- Using an Enterprise Router to Enable Network Communications for the Same Region but Different Tenants
- Using a Cloud Connection to Enable Cross-Region Network Communications
- Creating a Real-Time Migration Job
- Configuring a Real-Time Migration Job
- Real-Time Migration Job O&M
- Field Type Mapping
-
Job Performance Optimization
- Overview
- Optimizing Job Parameters
- Optimizing the Parameters of a Job for Migrating Data from MySQL to MRS Hudi
- Optimizing the Parameters of a Job for Migrating Data from MySQL to GaussDB(DWS)
- Optimizing the Parameters of a Job for Migrating Data from MySQL to DMS for Kafka
- Optimizing the Parameters of a Job for Migrating Data from DMS for Kafka to OBS
- Optimizing the Parameters of a Job for Migrating Data from Apache Kafka to MRS Kafka
- Optimizing the Parameters of a Job for Migrating Data from SQL Server to MRS Hudi
- Optimizing the Parameters of a Job for Migrating Data from PostgreSQL to GaussDB(DWS)
- Optimizing the Parameters of a Job for Migrating Data from Oracle to GaussDB(DWS)
- Optimizing the Parameters of a Job for Migrating Data from Oracle to MRS Hudi
-
Tutorials
- Overview
- Migrating a DRS Task to DataArts Migration
- Configuring a Job for Synchronizing Data from MySQL to MRS Hudi
- Configuring a Job for Synchronizing Data from MySQL to GaussDB(DWS)
- Configuring a Job for Synchronizing Data from MySQL to Kafka
- Configuring a Job for Synchronizing Data from DMS for Kafka to OBS
- Configuring a Job for Synchronizing Data from Apache Kafka to MRS Kafka
- Configuring a Job for Synchronizing Data from SQL Server to MRS Hudi
- Configuring a Job for Synchronizing Data from PostgreSQL to GaussDB(DWS)
- Configuring a Job for Synchronizing Data from Oracle to GaussDB(DWS)
- Configuring a Job for Synchronizing Data from Oracle to MRS Hudi
- Configuring a Job for Synchronizing Data from MongoDB to GaussDB(DWS)
- DataArts Architecture
-
DataArts Factory
- Overview
- Data Management
- Script Development
-
Job Development
- Job Development Process
- Creating a Job
- Developing a Pipeline Job
- Developing a Batch Processing Single-Task SQL Job
- Developing a Real-Time Processing Single-Task MRS Flink SQL Job
- Developing a Real-Time Processing Single-Task MRS Flink Jar Job
- Developing a Real-Time Processing Single-Task DLI Spark Job
- Setting Up Scheduling for a Job
- Submitting a Version
- Releasing a Job Task
- (Optional) Managing Jobs
- Notebook Development
- Solution
- Execution History
- O&M and Scheduling
- Configuration and Management
- Review Center
- Download Center
-
Node Reference
- Node Overview
- Node Lineages
- CDM Job
- Data Migration
- DIS Stream
- DIS Dump
- DIS Client
- Rest Client
- Import GES
- MRS Kafka
- Kafka Client
- ROMA FDI Job
- DLI Flink Job
- DLI SQL
- DLI Spark
- DWS SQL
- MRS Spark SQL
- MRS Hive SQL
- MRS Presto SQL
- MRS Spark
- MRS Spark Python
- MRS ClickHouse
- MRS Impala SQL
- MRS Flink Job
- MRS MapReduce
- CSS
- Shell
- RDS SQL
- ETL Job
- Python
- DORIS SQL
- ModelArts Train
- Create OBS
- Delete OBS
- OBS Manager
- Open/Close Resource
- Data Quality Monitor
- Subjob
- For Each
- SMN
- Dummy
- EL Expression Reference
- Simple Variable Set
-
Usage Guidance
- Referencing Parameters in Scripts and Jobs
- Setting the Job Scheduling Time to the Last Day of Each Month
- Configuring a Yearly Scheduled Job
- Using PatchData
- Obtaining the Output of an SQL Node
- Obtaining the Maximum Value and Transferring It to a CDM Job Using a Query SQL Statement
- IF Statements
- Obtaining the Return Value of a Rest Client Node
- Using For Each Nodes
- Using Script Templates and Parameter Templates
- Developing a Python Job
- Developing a DWS SQL Job
- Developing a Hive SQL Job
- Developing a DLI Spark Job
- Developing an MRS Flink Job
- Developing an MRS Spark Python Job
- DataArts Quality
- DataArts Catalog
-
DataArts Security
- Overview
- Dashboard
- Unified Permission Governance
- Sensitive Data Governance
- Sensitive Data Protection
- Data Security Operations
- Managing the Recycle Bin
-
DataArts DataService
- Overview
- Specifications
- Developing APIs in DataArts DataService
-
Calling APIs in DataArts DataService
- Applying for API Authorization
-
Calling APIs Using Different Methods
- API Calling Methods
- (Recommended) Using an SDK to Call an API Which Uses App Authentication
- Using an API Tool to Call an API Which Uses App Authentication
- Using an API Tool to Call an API Which Uses IAM Authentication
- Using an API Tool to Call an API Which Requires No Authentication
- Using a Browser to Call an API Which Requires No Authentication
- Viewing API Access Logs
- Configuring Review Center
- Audit Log
-
Best Practices
-
Advanced Data Migration Guidance
- Incremental Migration
- Using Macro Variables of Date and Time
- Migration in Transaction Mode
- Encryption and Decryption During File Migration
- MD5 Verification
- Configuring Field Converters
- Adding Fields
- Migrating Files with Specified Names
- Regular Expressions for Separating Semi-structured Text
- Recording the Time When Data Is Written to the Database
- File Formats
- Converting Unsupported Data Types
-
Advanced Data Development Guidance
- Dependency Policies for Periodic Scheduling
- Scheduling by Discrete Hours and Scheduling by the Nearest Job Instance
- Using PatchData
- Setting the Job Scheduling Time to the Last Day of Each Month
- Obtaining the Output of an SQL Node
- IF Statements
- Obtaining the Return Value of a Rest Client Node
- Using For Each Nodes
- Invoking DataArts Quality Operators Using DataArts Factory and Transferring Quality Parameters During Job Running
- Scheduling Jobs Across Workspaces
-
DataArts Studio Data Migration Configuration
- Overview
- Management Center Data Migration Configuration
- DataArts Migration Data Migration Configuration
- DataArts Architecture Data Migration Configuration
- DataArts Factory Data Migration Configuration
- DataArts Quality Data Migration Configuration
- DataArts Catalog Data Migration Configuration
- DataArts Security Data Migration Configuration
- DataArts DataService Data Migration Configuration
- Least Privilege Authorization
- How Do I View the Number of Table Rows and Database Size?
- Comparing Data Before and After Data Migration Using DataArts Quality
- Configuring Alarms for Jobs in DataArts Factory of DataArts Studio
- Scheduling a CDM Job by Transferring Parameters Using DataArts Factory
- Enabling Incremental Data Migration Through DataArts Factory
- Creating Table Migration Jobs in Batches Using CDM Nodes
- Automatic Construction and Analysis of Graph Data
- Simplified Migration of Trade Data to the Cloud and Analysis
- Migration of IoV Big Data to the Lake Without Loss
- Real-Time Alarm Platform Construction
-
Advanced Data Migration Guidance
- SDK Reference
-
API Reference
- Before You Start
- API Overview
- Calling APIs
-
DataArts Migration APIs
-
Cluster Management
- Querying Cluster Details
- Deleting a Cluster
- Querying All AZs
- Querying Supported Versions
- Querying Version Specifications
- Querying Details About a Flavor
- Querying the Enterprise Project IDs of All Clusters
- Querying the Enterprise Project ID of a Specified Cluster
- Query a Specified Instance in a Cluster
- Modifying a Cluster
- Restarting a Cluster
- Starting a Cluster
- Stopping a Cluster (To Be Taken Offline)
- Creating a Cluster
- Querying the Cluster List
- Job Management
- Link Management
- Public Data Structures
-
Cluster Management
-
DataArts Factory APIs (V1)
- Script Development APIs
- Resource Management APIs
-
Job Development APIs
- Creating a Job
- Modifying a Job
- Viewing a Job List
- Viewing Job Details
- Viewing a Job File
- Exporting a Job
- Batch Exporting Jobs
- Importing a Job
- Executing a Job Immediately
- Starting a Job
- Stopping a Job
- Deleting a Job
- Stopping a Job Instance
- Rerunning a Job Instance
- Viewing Running Status of a Real-Time Job
- Viewing a Job Instance List
- Viewing Job Instance Details
- Querying System Task Details
-
Connection Management APIs (To Be Taken Offline)
- Creating a Connection (to Be Taken Offline)
- Querying a Connection List (to Be Taken Offline)
- Querying Connection Details (to Be Taken Offline)
- Modifying a Connection (to Be Taken Offline)
- Deleting a Connection (to Be Taken Offline)
- Exporting Connections (to Be Taken Offline)
- Importing Connections (to Be Taken Offline)
-
DataArts Factory APIs (V2)
-
Job Development APIs
- Creating a PatchData Instance
- Querying PatchData Instances
- Stopping a PatchData Instance
- Changing a Job Name
- Querying Release Packages
- Querying Details About a Release Package
- Configuring Job Tags
- Querying Alarm Notifications
- Releasing Task Packages
- Canceling Task Packages
- Querying the Instance Execution Status
- Querying Completed Tasks
- Querying Instances of a Specified Job
-
Job Development APIs
-
DataArts Architecture APIs
- Overview
- Information Architecture
- Data Standards
- Data Sources
- Process Architecture
- Data Standard Templates
- Approval Management
- Subject Management
- Subject Levels
- Catalog Management
- Atomic Metrics
- Derivative Metrics
- Compound Metrics
- Dimensions
- Filters
- Dimension Tables
- Fact Tables
- Summary Tables
- Business Metrics
- Version Information
-
ER Modeling
- Lookup Table Model List
- Creating a Table Model
- Updating a Table Model
- Deleting a Table Model
- Querying a Relationship
- Viewing Relationship Details
- Querying All Relationships in a Model
- Viewing Table Model Details
- Obtaining a Model
- Creating a Model Workspace
- Updating the Model Workspace
- Deleting a Model Workspace
- Viewing Details About a Model
- Querying Destination Tables and Fields (To Be Offline)
- Exporting DDL Statements of Tables in a Model
- Converting a Logical Model to a Physical Model
- Obtaining the Operation Result
- Import and Export
- Customized Items
- Quality Rules
- Tag API
- Lookup Table Management
- DataArts Quality APIs
-
DataArts DataService APIs
-
API Management
- Create an API
- Querying an API List
- Updating an API
- Querying API Information
- Deleting APIs
- Publishing an API
- API operations (offline/suspension/resumption)
- Batch Authorization API (Exclusive Edition)
- Debugging an API
- API authorization operations (authorization/authorization cancellation/application/renewal)
- Querying API Publishing Messages in DLM Exclusive
- Querying Instances for API Operations in DLM Exclusive
- Querying API Debugging Messages in DLM Exclusive
- Importing an Excel File Containing APIs
- Exporting an Excel File Containing APIs
- Exporting a .zip File Containing All APIs
- Downloading an Excel Template
- Application Management
- Message Management
- Authorization Management
-
Service Catalog Management
- Obtaining the List of APIs and Catalogs in a Catalog
- Obtaining the List of APIs in a Catalog
- Obtaining the List of Sub-Catalogs in a Catalog
- Updating a Service Catalog
- Query the service catalog
- Creating a Service Catalog
- Deleting Directories in Batches
- Moving a Catalog to Another Catalog
- Moving APIs to Another Catalog
- Obtaining the ID of a Catalog Through Its Path
- Obtaining the Path of a Catalog Through Its ID
- Obtaining the Paths to a Catalog Through Its ID
- Querying the Service Catalog API List
- Gateway Management
- App Management
-
Overview
- Querying and Collecting Statistics on User-related Overview Development Indicators
- This API is used to query and collect statistics on user-related overview invoking metrics.
- Querying Top N API Services Invoked
- Querying Top N Services Used by an App
- Querying API Statistics Details
- Querying App Statistics
- Querying API Dashboard Data Details
- Querying Data Details of a Specified API Dashboard
- Querying App Dashboard Data Details
- Querying Top N APIs Called by a Specified API Application
- Cluster Management
-
API Management
- Application Cases
- Appendix
-
FAQs
-
Consultation and Billing
- How Do I Select a Region and an AZ?
- What Is a Database, Data Warehouse, Data Lake, and Huawei FusionInsight Intelligent Data Lake? What Are the Differences and Relationships Between Them?
- What Is the Relationship Between DataArts Studio and Huawei Horizon Digital Platform?
- What Are the Differences Between DataArts Studio and ROMA?
- Can DataArts Studio Be Deployed in a Local Data Center or on a Private Cloud?
- How Do I Create a Fine-Grained Permission Policy in IAM?
- How Do I Isolate Workspaces So That Users Cannot View Unauthorized Workspaces?
- What Should I Do If a User Cannot View Workspaces After I Have Assigned the Required Policy to the User?
- What Should I Do If Insufficient Permissions Are Prompted When I Am Trying to Perform an Operation as an IAM User?
- Can I Delete DataArts Studio Workspaces?
- Can I Transfer a Purchased or Trial Instance to Another Account?
- Does DataArts Studio Support Version Upgrade?
- Does DataArts Studio Support Version Downgrade?
- How Do I View the DataArts Studio Instance Version?
- Why Can't I Select a Specified IAM Project When Purchasing a DataArts Studio Instance?
- What Is the Session Timeout Period of DataArts Studio? Can the Session Timeout Period Be Modified?
- Will My Data Be Retained If My Package Expires or My Pay-per-Use Resources Are in Arrears?
- How Do I Check the Remaining Validity Period of a Package?
- Why Isn't the CDM Cluster in a DataArts Studio Instance Billed?
- Why Does the System Display a Message Indicating that the Number of Daily Executed Nodes Has Reached the Upper Limit? What Should I Do?
-
Management Center
- Which Data Sources Can DataArts Studio Connect To?
- What Are the Precautions for Creating Data Connections?
- What Should I Do If Database or Table Information Cannot Be Obtained Through a GaussDB(DWS)/Hive/HBase Data Connection?
- Why Are MRS Hive/HBase Clusters Not Displayed on the Page for Creating Data Connections?
- What Should I Do If a GaussDB(DWS) Connection Test Fails When SSL Is Enabled for the Connection?
- Can I Create Multiple Connections to the Same Data Source in a Workspace?
- Should I Select the API or Proxy Connection Type When Creating a Data Connection in Management Center?
- How Do I Migrate the Data Development Jobs and Data Connections from One Workspace to Another?
-
DataArts Migration (CDM Jobs)
- What Are the Differences Between CDM and Other Data Migration Services?
- What Are the Advantages of CDM?
- What Are the Security Protection Mechanisms of CDM?
- How Do I Reduce the Cost of Using CDM?
- Will I Be Billed If My CDM Cluster Does Not Use the Data Transmission Function?
- Why Am I Billed Pay per Use When I Have Purchased a Yearly/Monthly CDM Incremental Package?
- How Do I Check the Remaining Validity Period of a Package?
- Can CDM Be Shared by Different Tenants?
- Can I Upgrade a CDM Cluster?
- How Is the Migration Performance of CDM?
- What Is the Number of Concurrent Jobs for Different CDM Cluster Versions?
- Does CDM Support Incremental Data Migration?
- Does CDM Support Field Conversion?
- What Component Versions Are Recommended for Migrating Hadoop Data Sources?
- What Data Formats Are Supported When the Data Source Is Hive?
- Can I Synchronize Jobs to Other Clusters?
- Can I Create Jobs in Batches?
- Can I Schedule Jobs in Batches?
- How Do I Back Up CDM Jobs?
- What Should I Do If Only Some Nodes in a HANA Cluster Can Communicate with the CDM Cluster?
- How Do I Use Java to Invoke CDM RESTful APIs to Create Data Migration Jobs?
- How Do I Connect the On-Premises Intranet or Third-Party Private Network to CDM?
- Does CDM Support Parameters or Variables?
- How Do I Set the Number of Concurrent Extractors for a CDM Migration Job?
- Does CDM Support Real-Time Migration of Dynamic Data?
- Can I Stop CDM Clusters?
- How Do I Obtain the Current Time Using an Expression?
- What Should I Do If the Log Prompts that the Date Format Fails to Be Parsed?
- What Can I Do If the Map Field Tab Page Cannot Display All Columns?
- How Do I Select Distribution Columns When Using CDM to Migrate Data to GaussDB(DWS)?
- What Do I Do If the Error Message "value too long for type character varying" Is Displayed When I Migrate Data to DWS?
- What Can I Do If Error Message "Unable to execute the SQL statement" Is Displayed When I Import Data from OBS to SQL Server?
- What Should I Do If the Cluster List Is Empty, I Have No Access Permission, or My Operation Is Denied?
- Why Is Error ORA-01555 Reported During Migration from Oracle to DWS?
- What Should I Do If the MongoDB Connection Migration Fails?
- What Should I Do If a Hive Migration Job Is Suspended for a Long Period of Time?
- What Should I Do If an Error Is Reported Because the Field Type Mapping Does Not Match During Data Migration Using CDM?
- What Should I Do If a JDBC Connection Timeout Error Is Reported During MySQL Migration?
- What Should I Do If a CDM Migration Job Fails After a Link from Hive to GaussDB(DWS) Is Created?
- How Do I Use CDM to Export MySQL Data to an SQL File and Upload the File to an OBS Bucket?
- What Should I Do If CDM Fails to Migrate Data from OBS to DLI?
- What Should I Do If a CDM Connector Reports the Error "Configuration Item [linkConfig.iamAuth] Does Not Exist"?
- What Should I Do If Error "Configuration Item [linkConfig.createBackendLinks] Does Not Exist" or "Configuration Item [throttlingConfig.concurrentSubJobs] Does Not Exist" Is Reported?
- What Should I Do If Message "CORE_0031:Connect time out. (Cdm.0523)" Is Displayed During the Creation of an MRS Hive Link?
- What Should I Do If Message "CDM Does Not Support Auto Creation of an Empty Table with No Column" Is Displayed When I Enable Auto Table Creation?
- What Should I Do If I Cannot Obtain the Schema Name When Creating an Oracle Relational Database Migration Job?
- What Should I Do If invalid input syntax for integer: "true" Is Displayed During MySQL Database Migration?
-
DataArts Migration (Real-Time Jobs)
- Overview
- How Do I Troubleshoot a Network Disconnection Between the Data Source and Resource Group?
- Which Ports Must Be Allowed by the Data Source Security Group So That DataArts Migration Can Access the Data Source?
- How Do I Configure a Spark Periodic Task for Hudi Compaction?
- What Should I Do If an Error Is Reported During DDL Synchronization of New Columns in a Real-Time MySQL-to-DWS Synchronization Job?
- Why Does DWS Filter the Null Value of the Primary Key During Real-Time Synchronization from MySQL to DWS?
- What Should I Do If a Job for Synchronizing Data from Kafka to DLI in Real Time Fails and "Array element access needs an index starting at 1 but was 0" Is Displayed?
- How Do I Grant the Log Archiving, Query, and Parsing Permissions of an Oracle Data Source?
- How Do I Manually Delete Replication Slots from a PostgreSQL Data Source?
-
DataArts Architecture
- What Is the Relationship Between Lookup Tables and Data Standards?
- What Are the Differences Between ER Modeling and Dimensional Modeling?
- What Data Modeling Methods Are Supported by DataArts Architecture?
- How Can I Use Standardized Data?
- Does DataArts Architecture Support Database Reversing?
- What Are the Differences Between the Metrics in DataArts Architecture and DataArts Quality?
- Why Doesn't the Table in the Database Change After I Have Modified Fields in an ER or Dimensional Model?
- Can I Configure Lifecycle Management for Tables?
- How Should I Select a Subject When a Public Dimension (Date, Region, Supplier, or Product) Is Shared by Multiple Subject Areas?
-
DataArts Factory
- How Many Jobs Can Be Created in DataArts Factory? Is There a Limit on the Number of Nodes in a Job?
- Does DataArts Studio Support Custom Python Scripts?
- How Can I Quickly Rectify a Deleted CDM Cluster Associated with a Job?
- Why Is There a Large Difference Between Job Execution Time and Start Time of a Job?
- Will Subsequent Jobs Be Affected If a Job Fails to Be Executed During Scheduling of Dependent Jobs? What Should I Do?
- What Should I Pay Attention to When Using DataArts Studio to Schedule Big Data Services?
- What Are the Differences and Relationships Between Environment Variables, Job Parameters, and Script Parameters?
- What Should I Do If a Job Log Cannot Be Opened and Error 404 Is Reported?
- What Should I Do If the Agency List Fails to Be Obtained During Agency Configuration?
- Why Can't I Select Specified Peripheral Resources When Creating a Data Connection in DataArts Factory?
- Why Can't I Receive Job Failure Alarm Notifications After I Have Configured SMN Notifications?
- Why Is There No Job Running Scheduling Log on the Monitor Instance Page After Periodic Scheduling Is Configured for a Job?
- Why Isn't the Error Cause Displayed on the Console When a Hive SQL or Spark SQL Scripts Fails?
- What Should I Do If the Token Is Invalid During the Execution of a Data Development Node?
- How Do I View Run Logs After a Job Is Tested?
- Why Does a Job Scheduled by Month Start Running Before the Job Scheduled by Day Is Complete?
- What Should I Do If Invalid Authentication Is Reported When I Run a DLI Script?
- Why Cannot I Select a Desired CDM Cluster in Proxy Mode When Creating a Data Connection?
- Why Is There No Job Running Scheduling Record After Daily Scheduling Is Configured for the Job?
- What Do I Do If No Content Is Displayed in Job Logs?
- Why Do I Fail to Establish a Dependency Between Two Jobs?
- What Should I Do If an Error Is Reported During Job Scheduling in DataArts Studio, Indicating that the Job Has Not Been Submitted?
- What Should I Do If an Error Is Reported During Job Scheduling in DataArts Studio, Indicating that the Script Associated with Node XXX in the Job Has Not Been Submitted?
- What Should I Do If a Job Fails to Be Executed After Being Submitted for Scheduling and an Error Displayed: Depend Job [XXX] Is Not Running Or Pause?
- How Do I Create Databases and Data Tables? Do Databases Correspond to Data Connections?
- Why Is No Result Displayed After a Hive Task Is Executed?
- Why Is the Last Instance Status On the Monitor Instance Page Either Successful or Failed?
- How Do I Configure Notifications for All Jobs?
- What Is the Maximum Number of Nodes That Can Be Executed Simultaneously?
- Can I Change the Time Zone of a DataArts Studio Instance?
- How Do I Synchronize the Changed Names of CDM Jobs to DataArts Factory?
- Why Does the Execution of an RDS SQL Statement Fail and an Error Is Reported Indicating That hll Does Not Exist?
- What Should I Do If Error Message "The account has been locked" Is Displayed When I Am Creating a DWS Data Connection?
- What Should I Do If a Job Instance Is Canceled and Message "The node start execute failed, so the current node status is set to cancel." Is Displayed?
- What Should I Do If Error Message "Workspace does not exists" Is Displayed When I Call a DataArts Factory API?
- Why Don't the URL Parameters for Calling an API Take Effect in the Test Environment When the API Can Be Called Properly Using Postman?
- What Should I Do If Error Message "Agent need to be updated?" Is Displayed When I Run a Python Script?
- Why Is an Execution Failure Displayed for a Node in the Log When the Node Status Is Successful?
- What Should I Do If an Unknown Exception Occurs When I Call a DataArts Factory API?
- Why Is an Error Message Indicating an Invalid Resource Name Is Displayed When I Call a Resource Creation API?
- Why Does a PatchData Task Fail When All PatchData Job Instances Are Successful?
- Why Is a Table Unavailable When an Error Message Indicating that the Table Already Exists Is Displayed During Table Creation from a DWS Data Connection?
- What Should I Do If Error Message "The throttling threshold has been reached: policy user over ratelimit,limit:60,time:1 minute." Is Displayed When I Schedule an MRS Spark Job?
- What Should I Do If Error Message "UnicodeEncodeError: 'ascii' codec can't encode characters in position 63-64: ordinal not in range(128)" Is Displayed When I Run a Python Script?
- What Should I Do If an Error Message Is Displayed When I View Logs?
- What Should I Do If a Shell/Python Node Fails and Error "session is down" Is Reported?
- What Should I Do If a Parameter Value in a Request Header Contains More Than 512 Characters?
- What Should I Do If a Message Is Displayed Indicating that the ID Does Not Exist During the Execution of a DWS SQL Script?
- How Do I Check Which Jobs Invoke a CDM Job?
- What Should I Do If Error Message "The request parameter invalid" Is Displayed When I Use Python to Call the API for Executing Scripts?
- What Should I Do If the Default Queue of a New DLI SQL Script in DataArts Factory Has Been Deleted?
- Does the Event-based Scheduling Type in DataArts Factory Support Offline Kafka?
-
DataArts Quality
- What Are the Differences Between Quality Jobs and Comparison Jobs?
- How Can I Confirm that a Quality Job or Comparison Job Is Blocked?
- How Do I Manually Restart a Blocked Quality Job or Comparison Job?
- How Do I View Jobs Associated with a Quality Rule Template?
- What Should I Do If the System Displays a Message Indicating that I Do Not Have the MRS Permission to Perform a Quality Job?
- DataArts Catalog
-
DataArts Security
- Why Isn't Data Masked Based on a Specified Rule After a Data Masking Task Is Executed?
- What Should I Do If a Message Is Displayed Indicating that Necessary Request Parameters Are Missing When I Approve a GaussDB(DWS) Permission Application?
- What Should I Do If Error Message "FATAL: Invalid username/password,login denied" Is Displayed During the GaussDB(DWS) Connectivity Check When Fine-grained Authentication Is Enabled?
- What Should I Do If Error Message "Failed to obtain the database" Is Displayed When I Select a Database in DataArts Factory After Fine-grained Authentication Is Enabled?
- Why Does the System Display a Message Indicating Insufficient Permissions During Permission Synchronization to DLI?
-
DataArts DataService
- What Languages Do DataArts DataService SDKs Support?
- What Can I Do If the System Displays a Message Indicating that the Proxy Fails to Be Invoked During API Creation?
- What Should I Do If the Background Reports an Error When I Access the Test App Through the Data Service API and Set Related Parameters?
- How Many Times Can a Subdomain Name Be Accessed Using APIs Every Day?
- Can Operators Be Transferred When API Parameters Are Transferred?
- What Should I Do If No More APIs Can Be Created When the API Quota in the Workspace Is Used Up?
- How Can I Access APIs of DataArts DataService Exclusive from the Internet?
- How Can I Access APIs of DataArts DataService Exclusive Using Domain Names?
- What Should I Do If It Takes a Long Time to Obtain the Total Number of Data Records of a Table Through an API If the Table Contains a Large Amount of Data?
-
Consultation and Billing
-
More Documents
-
User Guide (Kuala Lumpur Region)
- Service Overview
- Preparations
-
User Guide
- Preparations Before Using DataArts Studio
- Management Center
-
DataArts Migration
- Overview
- Constraints
- Supported Data Sources
- Managing Clusters
-
Managing Links
- Creating Links
- Managing Drivers
- Managing Agents
- Managing Cluster Configurations
- Link to a Common Relational Database
- Link to a Database Shard
- Link to MyCAT
- Link to a Dameng Database
- Link to a MySQL Database
- Link to an Oracle Database
- Link to DLI
- Link to Hive
- Link to HBase
- Link to HDFS
- Link to OBS
- Link to an FTP or SFTP Server
- Link to Redis/DCS
- Link to DDS
- Link to CloudTable
- Link to CloudTable OpenTSDB
- Link to MongoDB
- Link to Cassandra
- Link to Kafka
- Link to DMS Kafka
- Link to Elasticsearch/CSS
- Managing Jobs
- Auditing
-
Tutorials
- Creating an MRS Hive Link
- Creating a MySQL Link
- Migrating Data from MySQL to MRS Hive
- Migrating Data from MySQL to OBS
- Migrating Data from MySQL to DWS
- Migrating an Entire MySQL Database to RDS
- Migrating Data from Oracle to CSS
- Migrating Data from Oracle to DWS
- Migrating Data from OBS to CSS
- Migrating Data from OBS to DLI
- Migrating Data from MRS HDFS to OBS
- Migrating the Entire Elasticsearch Database to CSS
- Advanced Operations
-
DataArts Factory
- Overview
- Data Management
- Script Development
- Job Development
- Solution
- Execution History
- O&M and Scheduling
- Configuration and Management
-
Node Reference
- Node Overview
- CDM Job
- Rest Client
- Import GES
- MRS Kafka
- Kafka Client
- ROMA FDI Job
- DLI Flink Job
- DLI SQL
- DLI Spark
- DWS SQL
- MRS Spark SQL
- MRS Hive SQL
- MRS Presto SQL
- MRS Spark
- MRS Spark Python
- MRS Flink Job
- MRS MapReduce
- CSS
- Shell
- RDS SQL
- ETL Job
- Python
- Create OBS
- Delete OBS
- OBS Manager
- Open/Close Resource
- Subjob
- For Each
- SMN
- Dummy
- EL Expression Reference
- Usage Guidance
-
FAQs
- Consultation
-
Management Center
- What Are the Precautions for Creating Data Connections?
- Why Do DWS/Hive/HBase Data Connections Fail to Obtain the Information About Database or Tables?
- Why Are MRS Hive/HBase Clusters Not Displayed on the Page for Creating Data Connections?
- What Should I Do If the Connection Test Fails When I Enable the SSL Connection During the Creation of a DWS Data Connection?
- Can I Create Multiple Data Connections in a Workspace in Proxy Mode?
- Should I Choose a Direct or a Proxy Connection When Creating a DWS Connection?
- How Do I Migrate the Data Development Jobs and Data Connections from One Workspace to Another?
- Can I Delete Workspaces?
-
DataArts Migration
- General
-
Functions
- Does CDM Support Incremental Data Migration?
- Does CDM Support Field Conversion?
- What Component Versions Are Recommended for Migrating Hadoop Data Sources?
- What Data Formats Are Supported When the Data Source Is Hive?
- Can I Synchronize Jobs to Other Clusters?
- Can I Create Jobs in Batches?
- Can I Schedule Jobs in Batches?
- How Do I Back Up CDM Jobs?
- How Do I Configure the Connection If Only Some Nodes in the HANA Cluster Can Communicate with the CDM Cluster?
- How Do I Use Java to Invoke CDM RESTful APIs to Create Data Migration Jobs?
- How Do I Connect the On-Premises Intranet or Third-Party Private Network to CDM?
- How Do I Set the Number of Concurrent Extractors for a CDM Migration Job?
- Does CDM Support Real-Time Migration of Dynamic Data?
-
Troubleshooting
- What Can I Do If Error Message "Unable to execute the SQL statement" Is Displayed When I Import Data from OBS to SQL Server?
- Why Is Error ORA-01555 Reported During Migration from Oracle to DWS?
- What Should I Do If the MongoDB Connection Migration Fails?
- What Should I Do If a Hive Migration Job Is Suspended for a Long Period of Time?
- What Should I Do If an Error Is Reported Because the Field Type Mapping Does Not Match During Data Migration Using CDM?
- What Should I Do If a JDBC Connection Timeout Error Is Reported During MySQL Migration?
- What Should I Do If a CDM Migration Job Fails After a Link from Hive to DWS Is Created?
- How Do I Use CDM to Export MySQL Data to an SQL File and Upload the File to an OBS Bucket?
- What Should I Do If CDM Fails to Migrate Data from OBS to DLI?
- What Should I Do If a CDM Connector Reports the Error "Configuration Item [linkConfig.iamAuth] Does Not Exist"?
- What Should I Do If Error Message "Configuration Item [linkConfig.createBackendLinks] Does Not Exist" Is Displayed During Data Link Creation or Error Message "Configuration Item [throttlingConfig.concurrentSubJobs] Does Not Exist" Is Displayed During Job Creation?
- What Should I Do If Message "CORE_0031:Connect time out. (Cdm.0523)" Is Displayed During the Creation of an MRS Hive Link?
- What Should I Do If Message "CDM Does Not Support Auto Creation of an Empty Table with No Column" Is Displayed When I Enable Auto Table Creation?
- What Should I Do If I Cannot Obtain the Schema Name When Creating an Oracle Relational Database Migration Job?
-
DataArts Factory
- How Many Jobs Can Be Created in DataArts Factory? Is There a Limit on the Number of Nodes in a Job?
- Why Is There a Large Difference Between Job Execution Time and Start Time of a Job?
- Will Subsequent Jobs Be Affected If a Job Fails to Be Executed During Scheduling of Dependent Jobs? What Should I Do?
- What Should I Pay Attention to When Using DataArts Studio to Schedule Big Data Services?
- What Are the Differences and Connections Among Environment Variables, Job Parameters, and Script Parameters?
- What Do I Do If Node Error Logs Cannot Be Viewed When a Job Fails?
- What Should I Do If the Agency List Fails to Be Obtained During Agency Configuration?
- How Do I Locate Job Scheduling Nodes with a Large Number?
- Why Cannot Specified Peripheral Resources Be Selected When a Data Connection Is Created in Data Development?
- Why Is There No Job Running Scheduling Log on the Monitor Instance Page After Periodic Scheduling Is Configured for a Job?
- Why Does the GUI Display Only the Failure Result but Not the Specific Error Cause After Hive SQL and Spark SQL Scripts Fail to Be Executed?
- What Do I Do If the Token Is Invalid During the Running of a Data Development Node?
- How Do I View Run Logs After a Job Is Tested?
- Why Does a Job Scheduled by Month Start Running Before the Job Scheduled by Day Is Complete?
- What Should I Do If Invalid Authentication Is Reported When I Run a DLI Script?
- Why Cannot I Select the Desired CDM Cluster in Proxy Mode When Creating a Data Connection?
- Why Is There No Job Running Scheduling Record After Daily Scheduling Is Configured for the Job?
- What Do I Do If No Content Is Displayed in Job Logs?
- Why Do I Fail to Establish a Dependency Between Two Jobs?
- What Should I Do If an Error Is Displayed During DataArts Studio Scheduling: The Job Does Not Have a Submitted Version?
- What Do I Do If an Error Is Displayed During DataArts Studio Scheduling: The Script Associated with Node XXX in the Job Is Not Submitted?
- What Should I Do If a Job Fails to Be Executed After Being Submitted for Scheduling and an Error Displayed: Depend Job [XXX] Is Not Running Or Pause?
- How Do I Create a Database And Data Table? Is the database a data connection?
- Why Is No Result Displayed After an HIVE Task Is Executed?
- Why Does the Last Instance Status On the Monitor Instance page Only Display Succeeded or Failed?
- How Do I Create a Notification for All Jobs?
- How Many Nodes Can Be Executed Concurrently in Each DataArts Studio Version?
- What Is the Priority of the Startup User, Execution User, Workspace Agency, and Job Agency?
-
API Reference (Kuala Lumpur Region)
- Before You Start
- API Overview
- Calling APIs
- Application Cases
-
DataArts Migration APIs
- Cluster Management
- Job Management
- Link Management
-
Public Data Structures
-
Link Parameter Description
- Link to a Relational Database
- Link to OBS
- Link to HDFS
- Link to HBase
- Link to CloudTable
- Link to Hive
- Link to an FTP or SFTP Server
- Link to MongoDB
- Link to Redis/DCS (to Be Brought Offline)
- Link to Kafka
- Link to Elasticsearch/Cloud Search Service
- Link to DLI
- Link to CloudTable OpenTSDB
- Link to Amazon S3
- Link to DMS Kafka
-
Source Job Parameters
- From a Relational Database
- From Object Storage
- From HDFS
- From Hive
- From HBase/CloudTable
- From FTP/SFTP/NAS (to Be Brought Offline)/SFS (to Be Brought Offline)
- From HTTP/HTTPS
- From MongoDB/DDS
- From Redis/DCS (to Be Brought Offline)
- From DIS
- From Kafka
- From Elasticsearch/Cloud Search Service
- From OpenTSDB
- Destination Job Parameters
- Job Parameter Description
-
Link Parameter Description
-
DataArts Factory APIs
- Connection Management APIs
- Script Development APIs
- Resource Management APIs
- Job Development APIs
- Data Structure
-
APIs to Be Taken Offline
- Creating a Job
- Editing a Job
- Viewing a Job List
- Viewing Job Details
- Exporting a Job
- Batch Exporting Jobs
- Importing a Job
- Executing a Job Immediately
- Starting a Job
- Viewing Running Status of a Real-Time Job
- Viewing a Job Instance List
- Viewing Job Instance Details
- Querying a System Task
- Creating a Script
- Modifying a Script
- Querying a Script
- Querying a Script List
- Querying the Execution Result of a Script Instance
- Creating a Resource
- Modifying a Resource
- Querying a Resource
- Querying a Resource List
- Importing a Connection
- Appendix
-
User Guide (Kuala Lumpur Region)
- General Reference
Copied.
Step 2: Prepare Data
Preparations Before Using DataArts Studio
If you are new to DataArts Studio, register a Huawei account, buy a DataArts Studio instance, create workspaces, and make other preparations. For details, see Buying and Configuring a DataArts Studio Instance. Then you can go to the created workspace and start using DataArts Studio.
In this example, the a Huawei account has all the permissions required for performing all the data operations on DataArts Studio so that the entire data governance process using DataArts Studio can be demonstrated.
Preparing a Data Source
This guide uses the collection of operations statistics from a taxi vendor in 2017 as an example.
The raw data of this example is from NYC open data platform.
You do not need to obtain the raw data. This example provides sample data that simulates the raw data. You can use the following method to prepare example data: Store example data in a .csv file, upload the .csv file to OBS, and use DataArts Migration of DataArts Studio to integrate the example data into other cloud services.
To prepare example data, perform the following steps:
- Create a CSV file (UTF-8 without BOM) named 2017_Yellow_Taxi_Trip_Data.csv, copy the sample data provided in the subsequent section to the CSV file, and save the file.
To generate a CSV file in Windows, you can perform the following steps:
- Use a text editor (for example, Notepad) to create a .txt document and copy the sample data to the document. Then check the total number of rows and check whether the data of rows is correctly separated. (If the sample data is copied from a PDF document, the data in a single row will be wrapped if the data is too long. In this case, you must manually adjust the data to ensure that it is in a single row.)
- Choose File > Save as. In the displayed dialog box, set Save as type to All files (*.*), enter the file name with the .csv suffix for File name, and select the UTF-8 encoding format (without BOM) to save the file in CSV format.
- Upload the CSV file to OBS.
- Log in to the management console and choose Storage > Object Storage Service to access the OBS console.
- Click Create Bucket and set parameters as prompted to create an OBS bucket named fast-demo.
NOTE:
To ensure network connectivity, select the same region for OBS bucket as that for the DataArts Studio instance. If an enterprise project is required, select the enterprise project that is the same as that of the DataArts Studio instance.
For details about how to create a bucket on the OBS console, see Creating a Bucket in Object Storage Service Console Operation Guide.
- Upload data to OBS bucket fast-demo.
For details about how to upload a file on the OBS console, see Uploading a File in Object Storage Service Console Operation Guide.
The example data is as follows:
VendorID,tpep_pickup_datetime,tpep_dropoff_datetime,passenger_count,trip_distance,RatecodeID,store_and_fwd_flag,PULocationID,DOLocationID,payment_type,fare_amount,extra,mta_tax,tip_amount,tolls_amount,improvement_surcharge,total_amount 2,02/14/2017 04:08:11 PM,02/14/2017 04:21:53 PM,1,0.91,1,N,237,163,2,9.5,1,0.5,0,0,0.3,11.3 2,02/14/2017 04:08:11 PM,02/14/2017 04:19:29 PM,2,1.03,1,N,237,229,1,8.5,1,0.5,2.06,0,0.3,12.36 1,02/14/2017 04:08:12 PM,02/14/2017 04:19:44 PM,1,1.6,1,N,186,163,2,9,1,0.5,0,0,0.3,10.8 1,02/14/2017 04:08:12 PM,02/14/2017 04:19:15 PM,1,1.2,1,N,48,48,2,8.5,1,0.5,0,0,0.3,10.3 2,02/14/2017 04:08:12 PM,02/14/2017 04:13:38 PM,5,0.61,1,N,161,162,1,5.5,1,0.5,2.19,0,0.3,9.49 2,02/14/2017 04:08:12 PM,02/14/2017 05:35:11 PM,1,19.31,2,N,152,132,1,52,4.5,0.5,12.57,5.54,0.3,75.41 1,02/14/2017 04:08:13 PM,02/14/2017 04:20:53 PM,1,1.9,1,N,236,143,1,10.5,1,0.5,1.85,0,0.3,14.15 2,02/14/2017 04:08:13 PM,02/14/2017 04:15:54 PM,1,0.61,1,N,48,164,1,6.5,1,0.5,1.66,0,0.3,9.96 2,02/14/2017 04:08:13 PM,02/14/2017 04:41:40 PM,1,6.04,1,N,244,262,1,25,1,0.5,6.7,0,0.3,33.5 2,02/14/2017 04:08:13 PM,02/14/2017 04:17:31 PM,1,1.39,1,N,170,234,1,8,1,0.5,1,0,0.3,10.8 2,02/14/2017 04:08:14 PM,02/14/2017 04:54:11 PM,2,10.12,1,N,140,189,1,37.5,1,0.5,7,0,0.3,46.3 2,02/14/2017 04:08:14 PM,02/14/2017 04:13:56 PM,1,0.71,1,N,179,7,2,5.5,1,0.5,0,0,0.3,7.3 2,02/14/2017 04:08:14 PM,02/14/2017 05:04:24 PM,1,18.1,2,N,263,132,1,52,4.5,0.5,15.71,5.54,0.3,78.55 2,02/14/2017 04:08:14 PM,02/14/2017 04:08:47 PM,1,0.02,1,N,231,231,2,2.5,1,0.5,0,0,0.3,4.3 2,02/14/2017 04:08:15 PM,02/14/2017 04:18:13 PM,1,1.34,1,N,100,162,1,8,1,0.5,1.2,0,0.3,11 1,02/14/2017 04:08:16 PM,02/14/2017 04:19:01 PM,1,1.8,1,N,239,151,1,9,1,0.5,2.15,0,0.3,12.95 2,02/14/2017 04:08:16 PM,02/14/2017 04:15:57 PM,1,1.06,1,N,68,170,1,6.5,1,0.5,1,0,0.3,9.3 2,02/14/2017 04:08:16 PM,02/14/2017 04:20:08 PM,2,1.5,1,N,161,142,1,9,1,0.5,2.16,0,0.3,12.96 2,02/14/2017 04:08:16 PM,02/14/2017 04:11:56 PM,1,0.62,1,N,87,88,2,4.5,1,0.5,0,0,0.3,6.3 2,02/14/2017 04:08:16 PM,02/14/2017 04:13:20 PM,1,0.88,1,N,262,236,2,5.5,1,0.5,0,0,0.3,7.3
No. |
Field Name |
Field Description |
---|---|---|
1 |
VendorID |
Vendor ID. Possible values are: 1=A Company 2=B Company |
2 |
tpep_pickup_datetime |
Time when a passenger gets on a taxi. |
3 |
tpep_dropoff_datetime |
Time when a passenger gets off a taxi. |
4 |
passenger_count |
Number of passengers. |
5 |
trip_distance |
Driving distance. |
6 |
ratecodeid |
Charge rate code. Possible values are: 1=Standard rate 2=JFK 3=Newark 4=Nassau or Westchester 5=Negotiated fare 6=Group ride |
7 |
store_fwd_flag |
Store-and-forward flag. |
8 |
PULocationID |
Location at which a passenger gets on a taxi. |
9 |
DOLocationID |
Location at which a passenger gets off a taxi. |
10 |
payment_type |
Payment type. Possible values are: 1=Credit card 2=Cash 3=No charge 4=Dispute 5=Unknown 6=Voided trip |
11 |
fare_amount |
Fare amount. |
12 |
extra |
Extra fee. |
13 |
mta_tax |
MTA tax. |
14 |
tip_amount |
Tip amount. |
15 |
tolls_amount |
Toll amount. |
16 |
improvement_surcharge |
Improvement surcharge. |
17 |
total_amount |
Total amount. |
Preparing a Data Lake
Before using DataArts Studio, you need to select cloud services or databases as the data foundation, which provides storage and compute capabilities. DataArts Studio provides one-stop data development, governance, and services based on the data foundation.
DataArts Studio can integrate cloud services such as GaussDB(DWS), DLI, and MRS Hive, as well as conventional databases such as MySQLOracle. For details, see Data Sources.
In this example, MapReduce Service (MRS) Hive is used as the data foundation of DataArts Studio. You need to create an MRS security cluster (that is, an MRS cluster with Kerberos authentication enabled). For details, see Buying a Custom Cluster.
To ensure that the MRS cluster can communicate with the DataArts Studio instance, the MRS cluster must meet the following requirements:
- The MRS cluster must contain a Hive component.
- If you want to enable automatic generation of quality jobs based on the data standards in DataArts Studio DataArts Architecture, ensure that the MRS cluster version is 2.0.3 or later and that the cluster contains Hive and Spark components and at least four nodes. In this example, this function is required.
If the connection fails after you select a cluster, check whether the MRS cluster can communicate with the CDM instance which functions as the agent. They can communicate with each other in the following scenarios:
- If the CDM cluster in the DataArts Studio instance and the MRS cluster are in different regions, a public network or a dedicated connection is required. If the Internet is used for communication, ensure that an EIP has been bound to the CDM cluster, and the MRS cluster can access the Internet and the port has been enabled in the firewall rule.
- If the CDM cluster in the DataArts Studio instance and the MRS cluster are in the same region, VPC, subnet, and security group, they can communicate with each other by default. If they are in the same VPC but in different subnets or security groups, you must configure routing rules and security group rules. For details about how to configure routing rules, see Configuring Routing Rules. For details about how to configure security group rules, see Configuring Security Group Rules.
- The MRS cluster and the DataArts Studio workspace belong to the same enterprise project. If they do not, you can modify the enterprise project of the workspace.
NOTE:
If an agent is connected to multiple MRS clusters and one of the MRS clusters is deleted or abnormal, connections to the other MRS clusters will be affected. Therefore, you are advised to connect an agent to only one MRS cluster.
Creating a Data Connection on Management Center
After the data lake is prepared, create a data connection on Management Center to connect to the cloud service that functions as the data lake.
- Log in to the DataArts Studio console by following the instructions in Accessing the DataArts Studio Instance Console.
- On the DataArts Studio console, locate a workspace and click Management Center.
- On the displayed Manage Data Connections page, click Create Data Connection.
Figure 1 Creating a data connection
- In the dialog box displayed, set data connection parameters and click OK.
The following part describes how to create an MRS Hive connection. See Figure 2 for details.
- Data Connection Type: MRS Hive is selected by default.
- Name: Enter mrs_hive_link.
- Tag: Enter a new tag name or select an existing tag from the drop-down list box. This parameter is optional.
- Applicable Modules: Retain the default settings.
- Connection Type: Select Proxy connection.
- Manual: Select Cluster Name Mode. IP and Port are automatically set.
- MRS Cluster Name: Select an existing MRS cluster.
- KMS Key: Select a KMS key and use it to encrypt sensitive data. If no KMS key is available, click Access KMS to go to the KMS console and create one.
- Agent: Select a DataArts Migration cluster as the connection agent. The DataArts Migration cluster and MRS cluster must be in the same region, AZ, VPC, and subnet, and the security group rule must allow communication between the two clusters. In this example, select the DataArts Migration cluster that is automatically created during DataArts Studio instance creation.
To connect to an MRS 2.x cluster, select the DataArts Migration cluster of the 2.x version as the agent.
- Username: Enter the Kerberos authentication user. In an MRS policy, user admin is the default management user and cannot be used as the authentication user of the cluster that uses Kerberos authentication. Therefore, to create a connection for an MRS cluster that uses Kerberos authentication, perform the following operations:
- Log in to MRS Manager as user admin.
- Choose System > Permission > Security Policy > Password Policy. Click Add Password Policy and add a policy under which the password never expires.
- Set Password Policy Name to neverexp.
- Set Password Validity Period (Days) to 0, indicating that the password never expires.
- Set Password Expiration Notification (Days) to 0.
- Retain the default values for other parameters.
- Choose System > Permission > User. On the page displayed, click Create to add a dedicated user as the Kerberos authentication user and set the password policy to neverexp. Select the user group superGroup for the user, and assign all roles to the user.
NOTE:
- For clusters of MRS 3.1.0 or later, the user must at least have permissions of the Manager_viewer role to create data connections in Management Center. To perform database, table, and data operations on components, the user must also have user group permissions of the components.
- For clusters earlier than MRS 3.1.0, the user must have permissions of the Manager_administrator or System_administrator role to create data connections in Management Center.
- A user with only the Manager_tenant or Manager_auditor permission cannot create connections.
- Log in to Manager as the new user and change the initial password. Otherwise, the connection fails to be created.
- Synchronize IAM users.
- Log in to the MRS console.
- Choose Clusters > Active Clusters, select a running cluster, and click its name to go to its details page.
- In the Basic Information area of the Dashboard page, click Synchronize on the right side of IAM User Sync to synchronize IAM users.
NOTE:
- When the policy of the user group to which the IAM user belongs changes from MRS ReadOnlyAccess to MRS CommonOperations, MRS FullAccess, or MRS Administrator, wait for 5 minutes until the new policy takes effect after the synchronization is complete because the SSSD (System Security Services Daemon) cache of cluster nodes needs time to be updated. Then, submit a job. Otherwise, the job may fail to be submitted.
- When the policy of the user group to which the IAM user belongs changes from MRS CommonOperations, MRS FullAccess, or MRS Administrator to MRS ReadOnlyAccess, wait for 5 minutes until the new policy takes effect after the synchronization is complete because the SSSD cache of cluster nodes needs time to be updated.
- Password: Enter the password of the Kerberos authentication user.
Creating a Database
According to the implementation process of data lake governance, you are advised to create a database for each of the layers (SDI layer, DWI layer, DWR layer, and DM layer) in the data lake to implement hierarchical sharding. Data sharding is a concept involved in DataArts Architecture.
- Source Data Integration (SDI) copies data from the source system.
- Data Warehouse Integration (DWI) integrates and cleanses data from multiple source systems, and builds ER models based on the third normal form (3NF).
- Data Warehouse Report (DWR) is based on the multi-dimensional model and its data granularity is the same as that of the DWI layer.
- Data Mart (DM) is where multiple types of data are summarized and displayed.
Generally, create a database in the data lake service.
In this example, you can use either of the following methods to create a database in MRS Hive:
- You can create a database on the DataArts Factory module of DataArts Studio. For details, see Creating a Database.
- You can also develop and execute a SQL script for creating a database using the DataArts Studio DataArts Factory module or on the MRS client, and then use the script to create a database. For details about how to develop a script in DataArts Factory, see Developing an SQL Script. For details about how to develop a script using the MRS Client, see Using Hive from Scratch. Run the following Hive SQL commands to create a database:
-- Create an SDI layer database. CREATE DATABASE demo_sdi_db; -- Create a DWI layer database. CREATE DATABASE demo_dwi_db; -- Create a DWR layer database. CREATE DATABASE demo_dwr_db; -- Create a DM layer database. CREATE DATABASE demo_dm_db;
Creating Tables
Based on sample data, create a source table to store raw data. To migrate data from a file to a database, you must create a destination table in advance. In this example, the data source is a CSV file on OBS instead of a database. When you use DataArts Studio DataArts Migration to migrate data to the cloud, the destination table cannot be automatically created. Therefore, you must create a table on the destination (MRS).
During data migration using DataArts Studio, a destination table can be automatically created for migration from relational databases to Hive and between relational databases. In this case, you do not need to create a table in the destination database in advance.
Run the following SQL statements to create a source table in the demo_sdi_db database to store raw data.
In this example, you can use either of the following methods to create a data table in MRS Hive:
- You can create a table on the DataArts Studio DataArts Factory module. For details, see Creating a Table.
- You can also develop and execute a SQL script for creating a table using the DataArts Studio DataArts Factory module or on the MRS client, and then use the script to create a table. For details about how to develop a script in DataArts Factory, see Developing an SQL Script. For details about how to develop a script using the MRS Client, see Using Hive from Scratch. The following is an example Hive SQL command used to create a raw table in the demo_sdi_db database.
DROP TABLE IF EXISTS `sdi_taxi_trip_data`; CREATE TABLE demo_sdi_db.`sdi_taxi_trip_data` ( `VendorID` BIGINT COMMENT '', `tpep_pickup_datetime` TIMESTAMP COMMENT '', `tpep_dropoff_datetime` TIMESTAMP COMMENT '', `passenger_count` BIGINT COMMENT '', `trip_distance` DECIMAL(10,2) COMMENT '', `ratecodeid` BIGINT COMMENT '', `store_fwd_flag` STRING COMMENT '', `PULocationID` STRING COMMENT '', `DOLocationID` STRING COMMENT '', `payment_type` BIGINT COMMENT '', `fare_amount` DECIMAL(10,2) COMMENT '', `extra` DECIMAL(10,2) COMMENT '', `mta_tax` DECIMAL(10,2) COMMENT '', `tip_amount` DECIMAL(10,2) COMMENT '', `tolls_amount` DECIMAL(10,2) COMMENT '', `improvement_surcharge` DECIMAL(10,2) COMMENT '', `total_amount` DECIMAL(10,2) COMMENT '' );
Feedback
Was this page helpful?
Provide feedbackThank you very much for your feedback. We will continue working to improve the documentation.See the reply and handling status in My Cloud VOC.
For any further questions, feel free to contact us through the chatbot.
Chatbot