Big Data Cloud Service Decision Guide
Overview
In digital transformation, data is the core asset of enterprises. As the key methods for mining data value, big data technologies are evolving from traditional centralized batch processing to real-time stream processing, interactive analysis, intelligent decision-making, and more.
Huawei Cloud big data services provide one-stop data processing, analytics, and governance for you to quickly gain insights from massive amounts of data and achieve data-driven business innovation.
Huawei Cloud big data services mainly include MapReduce Service (MRS), Data Warehouse Service (DWS), DataArts Studio, Cloud Search Service (CSS), and LakeFormation. These cloud services work together as a comprehensive big data technology stack on Huawei Cloud, providing data collection, storage, processing, analytics, and search capabilities.
Service layers
- DataArts Studio is a unified big data development and governance platform. It provides end-to-end data governance capabilities, including data integration, development, quality, and assets. It connects to other big data services and serves as a coordination hub.
- MRS, DWS, and CSS provide data processing capabilities for different scenarios.
- LakeFormation provides unified management of engine metadata, permissions, and transactions. This enables free flow of data and models in the data lake and seamless integration of the lake, warehouse, and AI.
Overview and Positioning of Core Services
DataArts Studio
DataArts Studio is a one-stop data governance and operations platform that provides intelligent data lifecycle management to help enterprises go digital. It provides end-to-end data governance capabilities, including data integration, development, architecture, quality, map, service, and security. These capabilities help enterprises quickly build an end-to-end intelligent data system from data ingestion to data analytics, eliminating data silos, unifying data standards, accelerating data monetization, and achieving digital transformation.
DataArts Studio provides the following core modules:
- DataArts Migration can migrate data between more than 20 data sources and ingest data into the data lake. It provides wizard-based configuration and management and supports incremental and periodic integration of data from a single table or an entire database.
- DataArts Factory enables you to quickly build a big data processing center, create data models, integrate data, develop scripts, and orchestrate workflows for data analytics.
- DataArts Architecture processes data so that it can be used for business. It enables you to plan data intelligently, customize data models, unify data standards, visualize data modeling, and label data to provide high-quality data for operational decision-making.
- DataArts Quality monitors data quality throughout data processing and generates real-time notifications for abnormal events.
- Data Map provides enterprise-level metadata management to clarify information assets. It visualizes data lineages and provides a panorama of data for intelligent data search, operations, and monitoring.
- DataArts DataService provides one-stop data service development, testing, and deployment. It enables agile response to data service needs, easier data acquisition, more efficient data consumption, and monetization of data assets.
- DataArts Security protects the security of data throughout the data lifecycle. It provides access permissions management, sensitive data identification, and privacy protection management to help you establish a security warning mechanism, enhance security protection, ensure data availability, and obtain security certifications.
MRS
Huawei Cloud MRS is a managed big data cluster service built on open-source Hadoop and Spark. It provides storage, computing, and analytics for massive amounts of data. MRS is compatible with APIs of open-source big data components. It supports O&M capabilities such as one-click cluster deployment, elastic scaling, automatic backup, and monitoring and alarm reporting, which greatly simplifies the deployment and O&M of Hadoop and Spark clusters.
MRS is suitable for big data scenarios such as offline data warehouses, real-time analytics, lakehouse, log analysis, and data mining.
MRS provides the following key features:
- One-click deployment: Clusters can be created in minutes. Multiple components and specifications are supported, including Hadoop, Spark, HBase, Kafka, and Flink.
- High-performance analysis: Huawei-developed Superior scheduler supports large clusters and responds to petabyte-level data queries in seconds.
- Lakehouse: Batch-stream convergence, storage-compute decoupling, and unified resource scheduling enable multi-dimensional analysis of data in a lake.
- Security and reliability: Kerberos authentication, role-based access control (RBAC), and table/column-level encryption ensure stability.
DWS
DWS is a cloud-based data warehousing solution that offers scalable, ready-to-use, and fully managed analytical database services. Adhering to ANSI/ISO SQL-92, SQL-99, and SQL:2003 standards, it is compatible with major database ecosystems like PostgreSQL, Oracle, Teradata, and MySQL. This makes it a competitive option for petabyte-scale big data analytics across diverse sectors.
DWS offers both storage-compute coupled and decoupled data warehouses and helps you create a cutting-edge data warehouse that excels in enterprise-level kernels, real-time analysis, collaborative computing, convergent analysis, and cloud native capabilities.
- Coupled storage and compute: This type of architecture provides enterprise-level data warehouse services with high performance, scalability, reliability, security, and easy O&M. It can analyze petabytes of data and is suitable for analytics that involve databases, warehouses, marts, and lakes.
- Decoupled storage and compute: This type of architecture allows compute and storage resources to scale independently and provides shared storage for virtual warehouses. These capabilities enable compute isolation and concurrent scaling to handle varied loads.
DWS can fit into a wide range of fields, such as finance, Internet of Vehicles (IoV), government and enterprise, e-commerce, energy, and telecommunications. It was listed in the Gartner Magic Quadrant for Data Management Solutions for two consecutive years, thanks to its large-scale scalability, enterprise-grade reliability, and higher cost-effectiveness over conventional data warehouses.
DWS provides the following key features:
- Ease of use: one-stop visualized management, seamless integration with other big data services, and migration of data between heterogeneous databases
- High performance: Parallel operator execution and vectorized execution engines enable query of trillions of data records in seconds.
- Scalability: online scale-out and upgrade, and elastic scaling of compute nodes
- High reliability: Distributed transactions support atomicity, consistency, isolation, durability (ACID) and strong data consistency. Logical components such as coordinator nodes (CNs) and data nodes (DNs) of clusters are designed with HA. Cluster-level and fine-grained backup and restoration, as well as cross-region disaster recovery, are supported.
- Low cost: DWS is billed on a pay-per-use basis and is available out of the box.
CSS
Huawei Cloud CSS is a managed search service built on open-source Elasticsearch, OpenSearch, and Logstash. It provides capabilities such as full-text search, log analysis, and security audit. CSS is compatible with Elasticsearch and OpenSearch APIs and supports O&M capabilities such as one-click cluster deployment, automatic backup, and monitoring and alarms. It significantly simplifies O&M of Elasticsearch and OpenSearch clusters. CSS is suitable for scenarios such as full-text search, log analysis, security audit, vector search, and semantic search.
CSS provides the following key features:
- Extensive ecosystem compatibility: CSS is fully compatible with native Elasticsearch and OpenSearch syntax, so there is no need for developers to change their habits. It supports mainstream client languages such as Python, Java, and Go, and can seamlessly integrate components such as Kibana, Dashboards, and Logstash.
- One-click deployment: Clusters can be created in minutes. Multiple specifications are supported.
- High-performance search: CSS supports inverted indexes and segmentation-based search with millisecond-level response times.
- Kernel enhancement: CSS offers a range of enhancements over the open-source kernel, such as decoupled storage and compute, large query isolation, and enhanced ingestion performance. These enhancements improve cluster performance and stability and reduce costs.
- Log analysis: CSS supports the ELK log analysis architecture and provides log collection, storage, search, and visualization.
- Vector database: Huawei-developed vector search engine enables high-performance, high-accuracy nearest neighbor or approximate nearest neighbor search over unstructured data, such as images, video, and language corpuses.
- Security compliance: CSS provides security capabilities such as user authentication, access control, and audit logs.
LakeFormation
Huawei Cloud LakeFormation is an enterprise-grade one-stop data lake construction and operations service with a storage-compute decoupled architecture. It provides a GUI and APIs for unified management of metadata in a data lake.
LakeFormation is compatible with the Hive metadata model and Ranger permission model, and can seamlessly work with Huawei Cloud big data services such as MRS and DWS. With LakeFormation, you can efficiently build and run data lakes to unlock data value.
LakeFormation is suitable for scenarios such as data lake construction, lakehouse analysis, cross-service data sharing, and unified metadata management.
LakeFormation provides the following key features:
- Unified metadata management: It can centrally manage metadata objects such as catalogs, databases, tables, and functions, enabling metadata sharing across engines.
- Fine-grained permission control: It is compatible with the Ranger permission model and supports database, table, and column permissions control. Granted permissions take effect across services, ensuring secure data sharing.
- Serverless high reliability: LakeFormation uses underlying resources to implement cross-AZ deployment, high reliability, auto scaling, unified metadata management, associated authorization of metadata and file directories, and interconnection with compute engines.
- Open ecosystem convergence: LakeFormation can seamlessly work with services such as MRS and DWS. It supports smooth migration of existing metadata and accelerates the implementation of lakehouse and data-AI convergence.
Key Decision-Making Dimensions
When selecting big data cloud services, you need to consider factors such as the service scenario, technical architecture, cost planning, and security compliance.
The following are the core dimensions for decision-making:
- Service scenario
- Data processing mode: Sort out core data processing requirements in advance and distinguish between scenarios such as batch processing, stream processing, interactive analysis, retrieval and search, and multimodal data processing. Batch processing is typically used for periodic offline computing, such as daily and monthly report generation. Stream processing is used for low-latency, real-time computing, such as real-time metric monitoring and personalized recommendation. Interactive analysis is used for instant self-service query, multi-dimensional BI dashboards, and self-service business exploration. Retrieval and search include fast full-text or multi-condition search of massive amounts of data, such as O&M log search, fuzzy product search, and user behavior search. Multimodal data processing is used for unified governance and computing of heterogeneous data. It supports converged analysis of structured tables, text, images, audio, videos, and time series IoT data.
- Data scale: Evaluate the data volume and predict growth.
- Timeliness requirements: Specify the timeliness requirements for data processing. For example, you can choose MRS Flink stream computing or DWS hybrid data warehouses for real-time processing, and choose MRS batch processing jobs for offline batch processing (hours or days).
- Technical architecture
- Decoupled and coupled storage and compute: The storage-compute decoupled architecture supports independent scale-out of storage and compute resources, provides higher resource utilization, and is suitable for elastic workloads. The storage-compute coupled architecture delivers better performance and is suitable for latency-sensitive workloads.
- Open-source compatibility and closed-source optimization: An open-source compatibility solution (such as MRS) facilitates migration and customization of service applications. A closed-source optimization solution provides better performance and stability.
- Semi-managed and fully-managed services: Semi-managed services allow you to control underlying capabilities of clusters to meet customized requirements. Fully-managed services eliminate the need for cluster O&M, reduce costs, and are suitable for fast deployment.
- Cost planning
- Billing mode: Choose the yearly/monthly or pay-per-use billing mode based on your workloads. Choose the yearly/monthly billing mode if your workloads are stable, and choose the pay-per-use or serverless billing mode if your workloads fluctuate.
- Resource utilization: Evaluate resource utilization and select a solution that can fully utilize resources. Decoupled storage and compute and serverless can significantly improve resource utilization.
- Data transmission costs: Consider the costs of data inflow and outflow and inter-cloud data transmission, and select a solution that allows for nearby data deployment.
- Security compliance
- Data security: Evaluate data sensitivity and select a solution that meets security requirements. Pay attention to security capabilities such as data encryption, access control, and audit logs.
- Compliance requirements: Select services based on industry compliance requirements.
- Data residency: Select a proper region and deployment mode based on data residency requirements.
Choosing Suitable Services
Huawei Cloud provides a comprehensive matrix of big data services, covering data storage, processing, analysis, governance, and search.
When selecting services, you are advised to:
- Determine the service scenario: Check whether your core requirement is data governance, real-time analysis, big data processing, or search.
- Evaluate the data volume: Select an appropriate service portfolio based on the data volume.
- Consider costs: Evaluate the characteristics of your workloads and select a suitable billing mode, such as yearly/monthly, pay-per-use, or serverless.
- Focus on security and compliance: Select services that meet data security and industry compliance requirements.
By selecting and combining Huawei Cloud big data services, you can quickly build a big data platform that meets your needs, unlocks the value of your data, and drives business innovation.
The following table compares the core services based on the preceding decision-making dimensions.
| Evaluation Dimension | MRS | DWS | DataArts Studio | CSS |
|---|---|---|---|---|
| Core positioning | Big data processing platform | High-performance data warehouse for OLAP | Data governance and development platform | Online distributed search service |
| Main data processing scenarios | Batch and stream processing | Real-time BI, analysis, and query | Data integration, development, and governance | Interactive analysis, full-text search, and log analysis |
| Data scale | Terabytes to petabytes (coupled storage and compute) | Terabytes to petabytes (coupled storage and compute) | This service does not store data. | Gigabytes to petabytes (coupled storage and compute) |
| Service data storage mode |
|
| This service does not store data. Instead, it integrates storage, databases, and big data services for data integration, development, and governance. | Local storage + OBS |
| Billing mode | Yearly/Monthly or pay-per-use | Yearly/Monthly or pay-per-use | Yearly/Monthly | Yearly/Monthly or pay-per-use |
| Typical scenarios | Data lake, machine learning, and real-time analytics | BI reports, real-time analytics, and real-time writing | Data integration, data platform, and data governance | Full-text search, log analysis, knowledge base Q&A, personalized recommendation, vector search, and semantic search |
| Typical Scenario | Recommended Service Portfolio | Description |
|---|---|---|
| Migration of a big data platform to the cloud | MRS + DataArts Studio | You can smoothly migrate workloads from your on-premises big data platform or other big data cloud service platforms to MRS, and quickly build an on-premises system based on the cloud environment to keep up with rapid business growth in the future. |
| Migration of a full-stack data warehouse platform to the cloud | DWS + DataArts Studio | To unlock the value of your data, you need to integrate data resources and build a big data platform. DWS Express can analyze the data that has been integrated and processed by a big data platform and stored in OBS. |
| Data platform building | MRS + DWS + DataArts Studio | To analyze and mine the value of massive amounts of data assets, you need to build a high-performance big data platform that supports data collection, storage, processing, analysis, and search. |
| Log analysis, O&M, and monitoring | CSS + Distributed Message Service (DMS) | CSS is used for analyzing ELB, server, container, and application logs, and DMS is used for ingesting data. |
Getting Started and Practices
You can quickly get started with Huawei Cloud big data services by following the instructions below.
- Beginners: DLI-powered Data Development Based on E-commerce BI Reports
Learn the functions of the DataArts Factory module, including how to edit scripts and jobs, and schedule jobs.
- Novices: DWS-powered Data Integration and Development Based on Movie Scores
Learn how to migrate data using the DataArts Migration module, and how to develop scripts, develop jobs, and schedule jobs using the DataArts Factory module.
- Experienced Users: MRS Hive-powered Data Governance Based on Taxi Trip Data
Learn how to perform end-to-end data operations with DataArts Studio.
- Comparing Data Before and After Data Migration Using DataArts Quality
Learn how to use the DataArts Quality module of DataArts Studio to check data consistency before and after data is migrated from DWS to an MRS Hive partitioned table.
- Configuring Alarms for Jobs in DataArts Factory of DataArts Studio
You can configure retries upon job node failures and failure alarms to minimize job failures during peak hours. Even if a job fails, O&M engineers can receive notifications and respond in a timely manner to prevent severer faults.
- Creating and Using a Hadoop Cluster for Offline Analysis
Learn how to create a Hadoop cluster for offline analysis and submit a wordcount job through the cluster client to count the number of words in big data.
- Creating and Using a Kafka Cluster for Stream Processing
Learn how to create a Kafka stream analysis cluster and generate and consume messages in a Kafka topic.
- Creating and Using a ClickHouse Cluster for Columnar Store
Learn how to create a ClickHouse cluster and create and query a ClickHouse table through the cluster client.
- Service Selection
Obtain the suggestions on selecting a suitable cluster type based on the service scenario, scale, performance, and cost.
- Migrating Data from Hadoop to MRS with CDM
Learn how to migrate Hadoop data to an MRS cluster using CDM. Based on the big data migration to the cloud and intelligent data lake solution, CDM provides easy-to-use migration capabilities and capabilities of integrating multiple data sources to the data lake, reducing the complexity of data source migration and integration and effectively improving the data migration and integration efficiency.
- Configuring Storage-Compute Decoupling for an MRS Cluster
MRS allows you to store service data in the OBS file system and use MRS clusters for data computing only. This storage-compute decoupled architecture enables you to scale resources on demand and analyze massive amounts of data at a low cost.
- Quickly Creating a DWS Cluster and Importing Data for Query
Learn how to create a DWS cluster, connect to a database, create a table, and perform a query.
- Getting Started with DWS Data Development SQL
Learn the basic SQL syntax.
- Selecting a Proper Table Type to Boost Development and Queries
Learn how to select a proper table type to improve development efficiency and query performance.
- DWS Performance Tuning
Learn the performance tuning methods for SQL statements, storage, and concurrency.
- Data Migration
Learn how to migrate data from Oracle, MySQL, ADB, and other databases to DWS.
- Using Elasticsearch for Data Search
Learn how to create a CSS cluster, create indexes, import data, and search for data.
- Using Elasticsearch for Vector Search
Learn how to create a vector database, store data, and search for vectors.
- Elasticsearch Data Migration
Learn how to migrate Elasticsearch data from different sources to CSS.
- Using Logstash to Sync Kafka Data to Elasticsearch
Learn how to use Logstash to efficiently synchronize Kafka data to Elasticsearch.
- Creating a DataArts Lake Formation Instance and Planning Metadata
Learn how to create a DataArts Lake Formation (LakeFormation) instance from scratch and set up catalogs along with internal databases, tables, and other metadata within the instance.
- Granting Catalog Operation Permissions to a LakeFormation Role
Learn how to create a LakeFormation role and grant it the permissions to modify catalogs and create databases.
- Configuring LakeFormation to Interconnect with Open-Source Spark
Learn how to interconnect LakeFormation with open-source Spark.
- Configuring LakeFormation to Interconnect with Open-Source Hive
Learn how to interconnect LakeFormation with open-source Hive.
- Configuring LakeFormation to Interconnect with MRS
Learn how to interconnect LakeFormation instances with MRS clusters to implement unified data lake metadata and permissions management.
What is your overall rating for this page?
Thank you very much for your feedback. We will continue working to improve the documentation.See the reply and handling status in My Cloud VOC.
For any further questions, feel free to contact us through the chatbot.
Chatbot