Help Center/ Analytics Decisions/ Decision Guides/ Big Data Cloud Service Decision Guide
Updated on 2026-09-20 GMT+08:00

Big Data Cloud Service Decision Guide

Overview

In digital transformation, data is a core enterprise asset. Big data technologies are key methods for mining data value, and they have been evolving from traditional centralized batch processing to real-time stream processing, interactive analysis, intelligent decision-making, and more.

Huawei Cloud big data services provide one-stop data processing, analytics, and governance that helps you quickly gain insights from massive amounts of data and achieve data-driven business innovation.

Huawei Cloud big data services mainly include MapReduce Service (MRS), Data Warehouse Service (DWS), DataArts Studio, Cloud Search Service (CSS), and LakeFormation. These cloud services work together to form a comprehensive big data technology stack on Huawei Cloud, providing data collection, storage, processing, analytics, and search capabilities.

Figure 1 Big data cloud services

Service layers

  • DataArts Studio is a unified big data development and governance platform. It provides end-to-end data governance capabilities, including data integration, development, quality, and assets. It connects to other big data services and serves as a coordination hub.
  • MRS, DWS, and CSS provide data processing capabilities for different scenarios.
  • LakeFormation provides unified management of engine metadata, permissions, and transactions. This enables free flow of data and models in the data lake and seamless integration of the lake, warehouse, and AI.

Overview and Positioning of Core Services

DataArts Studio

DataArts Studio is a one-stop data governance and operations platform that provides intelligent data lifecycle management to help enterprises go digital. It provides end-to-end data governance across data integration, development, architecture, quality, maps, services, and security, helping enterprises quickly build an end-to-end intelligent data system, covering everything from data ingestion to data analytics, eliminating data silos, unifying data standards, accelerating data monetization, and achieving digital transformation.

DataArts Studio provides the following components and core capabilities:

  • DataArts Migration can migrate data between more than 20 data sources and ingest data into a data lake. It provides wizard-based configuration and management and supports incremental and periodic integration of data from a single table or an entire database.
  • DataArts Factory provides an easy-to-use big data development environment that enables you to quickly build a big data processing center, create data models, integrate data, develop scripts, and orchestrate workflows for data analytics.
  • DataArts Architecture is a core module of DataArts Studio and processes data so that it can be used for business. It enables you to plan data intelligently, customize data models, unify data standards, visualize data modeling, and label data to provide high-quality data for operational decision-making.
  • DataArts Quality monitors data quality throughout the data processing process, with real-time notifications reported for abnormal events.
  • Data Map provides enterprise-level metadata management and provides clarity for information assets. It visualizes data lineages and provides a bird's-eye view of data for intelligent data search, operations, and monitoring.
  • DataArts DataService provides one-stop data service development, testing, and deployment. It enables agile response to data service needs, easier data acquisition, more efficient data consumption, and monetization of data assets.
  • DataArts Security protects the security of your data throughout its lifecycle. It provides access permissions management, sensitive data identification, and privacy protection management to help you establish a security warning mechanism, enhance security protection, and ensure data availability while protecting data security and compliance.

MRS

Huawei Cloud MRS is a managed big data cluster service built on open-source Hadoop and Spark. It provides storage, computing, and analytics for massive amounts of data. MRS is compatible with APIs of open-source big data components. It supports O&M capabilities such as one-click cluster deployment, elastic scaling, automatic backup, and monitoring and alarm reporting, which greatly simplifies the deployment and O&M of Hadoop and Spark clusters.

MRS is suitable for big data scenarios such as offline data warehouses, real-time analytics, lakehouse, and log analysis.

MRS provides the following key features:

  • One-click deployment: Clusters can be created in minutes. Multiple components and specifications are supported, including Hadoop, Spark, HBase, Kafka, Flink, and Hive.
  • High-performance analysis: Huawei-developed Superior scheduler supports large clusters and responds to petabyte-level data queries in seconds.
  • Lakehouse: Batch-stream convergence, storage-compute decoupling, and unified resource scheduling enable multi-dimensional analysis of data within a lake.
  • Security and reliability: Kerberos authentication, role-based access control (RBAC), and table/column-level encryption ensure stability.

DWS

DWS is a cloud-based data warehousing solution that offers scalable, ready-to-use, and fully managed analytical database services. Adhering to ANSI/ISO SQL-92, SQL-99, and SQL:2003 standards, it is compatible with major database ecosystems like PostgreSQL, Oracle, Teradata, and MySQL. This makes it a competitive option for petabyte-scale big data analytics across diverse sectors.

DWS offers data warehouses with both coupled and decoupled storage and compute to help you create a cutting-edge data warehouse that excels in enterprise-level kernels, real-time analysis, collaborative computing, convergent analysis, and cloud native capabilities.

  • Coupled storage and compute: This type of architecture provides enterprise-level data warehouse services with high performance, scalability, reliability, security, and easy O&M. It allows you to analyze petabytes of data and is suitable for analytics that involve databases, warehouses, marts, and lakes.
  • Decoupled storage and compute: This cloud native architecture allows compute and storage resources to scale independently and provides shared storage for virtual warehouses. These capabilities enable compute isolation and concurrent scaling to handle varied loads.

DWS can fit into a wide range of fields, such as finance, Internet of Vehicles (IoV), government and enterprise, e-commerce, energy, and telecommunications. It was listed in the Gartner Magic Quadrant for Data Management Solutions for two consecutive years, thanks to its large-scale scalability, enterprise-grade reliability, and significantly higher cost-effectiveness over conventional data warehouses.

DWS provides the following key features:

  • Ease of use: one-stop visualized management, seamless integration with other big data services, and migration of data between heterogeneous databases
  • High performance: Parallel operator execution and vectorized execution engines enable query of trillions of data records in seconds.
  • Scalability: online scale-out and upgrade, and elastic scaling of compute nodes
  • High reliability: Distributed transactions support atomicity, consistency, isolation, durability (ACID) and strong data consistency. Logical components such as coordinator nodes (CNs) and data nodes (DNs) of clusters are designed with HA. Cluster-level and fine-grained backup and restoration, as well as cross-region disaster recovery, are supported.
  • Low cost: DWS is billed on a pay-per-use basis and is ready to use out of the box.

CSS

Huawei Cloud CSS is a managed search service built on open-source Elasticsearch, OpenSearch, and Logstash. It provides capabilities such as full-text search, log analysis, and security audit. CSS is compatible with Elasticsearch and OpenSearch APIs and supports O&M capabilities such as one-click cluster deployment, automatic backup, and monitoring and alarms. It significantly simplifies O&M of Elasticsearch and OpenSearch clusters. CSS is suitable for scenarios such as full-text search, log analysis, security audit, vector search, and semantic search.

CSS provides the following key features:

  • Extensive ecosystem compatibility: CSS is fully compatible with native Elasticsearch and OpenSearch syntax, so there is no need for developers to change their habits. It supports software development kits (SDKs) of mainstream languages such as Python, Java, and Go, and can seamlessly integrate components such as Kibana, Dashboards, and Logstash.
  • One-click deployment: Clusters can be created in minutes. Multiple specifications are supported.
  • High-performance search: CSS supports inverted indexes and segmentation-based search with millisecond-level response times.
  • Kernel enhancement: CSS offers a range of enhancements over the open-source kernel, such as decoupled storage and compute, large query isolation, and enhanced ingestion performance. These enhancements improve cluster performance and stability and reduce costs.
  • Log analysis: CSS supports the ELK log analysis architecture and provides log collection, storage, search, and visualization.
  • Vector database: Huawei-developed vector search engine enables high-performance, high-accuracy nearest neighbor or approximate nearest neighbor search over unstructured data, such as images, video, and language corpuses.
  • Security compliance: CSS provides security capabilities such as user authentication, access control, and audit logs.

LakeFormation

Huawei Cloud LakeFormation is an enterprise-grade one-stop data lake construction and operations service with a storage-compute decoupled architecture. It provides a GUI and APIs for unified management of metadata in a data lake.

LakeFormation is compatible with the Hive metadata model and Ranger permission model, and can seamlessly work with Huawei Cloud big data services such as MRS and DWS. With LakeFormation, you can efficiently build and run data lakes to unlock data value.

LakeFormation is suitable for scenarios such as data lake construction, lakehouse analysis, cross-service data sharing, and unified metadata management.

LakeFormation provides the following key features:

  • Unified metadata management: It can centrally manage metadata objects such as catalogs, databases, tables, and functions, enabling metadata sharing across engines.
  • Fine-grained permission control: It is compatible with the Ranger permission model and supports database, table, and column permissions control. Granted permissions take effect across services, ensuring secure data sharing.
  • Serverless high reliability: LakeFormation uses underlying resources to implement cross-AZ deployment, high reliability, auto scaling, unified metadata management, associated authorization of metadata and file directories, and interconnection with compute engines.
  • Open ecosystem convergence: LakeFormation can seamlessly work with services such as MRS and DWS. It supports smooth migration of existing metadata and accelerates the implementation of lakehouse and data-AI convergence.

Key Decision-Making Dimensions

When selecting big data cloud services, you need to consider factors such as the service scenario, technical architecture, cost planning, and security compliance.

The following are the key factors you are advised to consider:

  • Service scenario
    • Data processing mode: Identify core data processing requirements in advance and distinguish between scenarios such as batch processing, stream processing, interactive analysis, retrieval and search, and multimodal data processing. Batch processing is typically used for periodic offline computing, such as daily and monthly report generation. Stream processing is used for low-latency, real-time computing, such as real-time metric monitoring and personalized recommendations. Interactive analysis is used for instant self-service query, multi-dimensional BI dashboards, and self-service business exploration. Retrieval and search include fast full-text or multi-condition search of massive amounts of data, such as O&M log search, fuzzy product search, and user behavior search. Multimodal data processing is used for unified governance and computing of heterogeneous data. It supports converged analysis of structured tables, text, images, audio, videos, and time series IoT data.
    • Data scale: Evaluate the data volume and predict growth.
    • Timeliness requirements: Specify the timeliness requirements for data processing. For example, you can choose MRS Flink stream computing or DWS hybrid data warehouses for real-time processing, and choose MRS batch processing jobs for offline batch processing (hours or days).
  • Technical architecture
    • Decoupled and coupled storage and compute: Decoupled storage and compute supports independent scale-out of storage and compute resources, provides higher resource utilization, and is suitable for elastic workloads. A coupled architecture delivers better performance and is suitable for latency-sensitive workloads.
    • Open-source compatibility and closed-source optimization: An open-source compatibility solution (such as MRS) facilitates migration and customization of service applications. A closed-source optimization solution provides better performance and stability.
    • Semi-managed and fully-managed services: Semi-managed services allow you to control underlying capabilities of clusters to meet customized requirements. Fully-managed services eliminate the need for cluster O&M, reduce costs, and are suitable for fast deployment.
  • Cost planning
    • Billing mode: Choose yearly/monthly or pay-per-use billing based on your workloads. Choose yearly/monthly billing if your workloads are stable, and choose pay-per-use billing if your workloads fluctuate.
    • Resource utilization: Evaluate resource utilization and select a solution that can most fully utilize resources. Decoupled storage and compute and serverless can significantly improve resource utilization.
    • Data transmission costs: Consider the costs of data inflow and outflow and inter-cloud data transmission, and select a solution that allows for nearby data deployment.
  • Security compliance
    • Data security: Evaluate data sensitivity and select a solution that meets security requirements. Consider security capabilities such as data encryption, access control, and audit logs.
    • Compliance requirements: Select services based on industry compliance requirements.
    • Data residency: Select a region and deployment mode based on data residency requirements.

Choosing Suitable Services

Huawei Cloud provides a comprehensive matrix of big data services, covering data storage, processing, analysis, governance, and search.

When selecting services, you are advised to:

  • Determine the service scenario: Check whether your core requirement is data governance, real-time analysis, big data processing, or search.
  • Evaluate the data volume: Select a service portfolio based on the data volume.
  • Consider costs: Evaluate the characteristics of your workloads and select yearly/monthly or pay-per-use billing.
  • Focus on security and compliance: Select services that meet data security and industry compliance requirements.

By selecting and combining Huawei Cloud big data services, you can quickly build a big data platform that meets your needs, unlocks the value of your data, and drives business innovation.

The following table compares the core services based on the preceding decision-making dimensions.

Table 1 Big data services

Evaluation Dimension

MRS

DWS

DataArts Studio

CSS

Core positioning

Big data processing platform

High-performance data warehouse for OLAP

Data governance and development platform

Online distributed search service

Main data processing scenarios

Batch processing, stream processing, and interactive analysis

Real-time BI, analysis and query, and ETL SQL batch processing

Data integration, development, and governance

Interactive analysis, full-text search, and log analysis

Data scale

Terabytes to petabytes (coupled storage and compute)

Terabytes to petabytes

This service does not store data.

Gigabytes to petabytes (coupled storage and compute)

Service data storage mode

  • Decoupled storage and compute: OBS
  • Coupled storage and compute: EVS
  • Decoupled storage and compute: OBS
  • Coupled storage and compute: EVS and local disks

This service does not store data. Instead, it connects to storage, databases, and big data services for data integration, development, and governance.

Local storage + OBS

Billing mode

Yearly/Monthly or pay-per-use

Yearly/Monthly or pay-per-use

Yearly/Monthly

Yearly/Monthly or pay-per-use

Typical scenarios

Data lake, machine learning, and real-time analytics

BI reports, real-time analytics, real-time writing, and ETL SQL batch processing

Data integration, data platform, and data governance

Full-text search, log analysis, knowledge base Q&A, personalized recommendation, vector search, and semantic search

Table 2 Recommended service portfolios for typical scenarios

Typical Scenario

Recommended Service Portfolio

Description

Migration of a big data platform to the cloud

MRS + DataArts Studio

You can smoothly migrate workloads from your on-premises big data platform or other big data cloud service platforms to MRS, and quickly build a system based on the cloud environment to keep up with rapid business growth in the future.

Migration of a full-stack data warehouse platform to the cloud

DWS + DataArts Studio

To unlock the value of your data, you need to integrate data resources and build a big data platform. DWS Express can analyze the data that has been integrated and processed by a big data platform and stored in OBS.

Data platform building

MRS + DWS + DataArts Studio

To analyze and mine the value of massive amounts of data assets, you need to build a high-performance big data platform that supports data collection, storage, processing, analysis, and search.

Log analysis and O&M monitoring

CSS + Distributed Message Service (DMS)

CSS is used for analyzing ELB, server, container, and application logs, and DMS is used for ingesting data.

Getting Started and Practices

You can quickly get started with Huawei Cloud big data services by following the instructions below.

  • Service Selection

    Obtain cluster type recommendations based on the service scenario, scale, performance, and cost.

  • Migrating Data from Hadoop to MRS with CDM

    Learn how to migrate Hadoop data to an MRS cluster using CDM. Based on the big data migration to the cloud and intelligent data lake solution, CDM provides easy-to-use migration capabilities and capabilities of integrating multiple data sources to the data lake, reducing the complexity of data source migration and integration and effectively improving the data migration and integration efficiency.

  • Configuring Storage-Compute Decoupling for an MRS Cluster

    MRS allows you to store service data in the OBS file system and use MRS clusters for data computing only. This storage-compute decoupled architecture enables you to scale resources on demand and analyze massive amounts of data at a low cost.