What Is MLOps
Time to read: 5 minutes

Machine Learning Operations (MLOps) is an engineering framework that standardizes the development, deployment, and O&M of machine learning models. By leveraging automated pipelines, MLOps ensures a seamless transition of models from experimental environments to production, while continuously monitoring model performance and automatically triggering retraining when necessary. ModelArts, a one-stop training and inference platform provided by Huawei Cloud, integrates built-in MLOps capabilities to help enterprises establish efficient AI engineering systems. These capabilities significantly shorten model rollout cycles and ensure stable, reliable model performance in production environments.

Why You Need MLOps

When promoting AI applications, enterprises usually rely on data scientists to train and validate models in experimental environments. However, significant gaps exist between these experimental settings and production environments. Code developed for experimentation is rarely ready for direct deployment, model performance can degrade online due shifts in data distribution, and manual model updates are inefficient and error-prone.

As AI applications grow, the proliferation of models has become exponential. Traditional manual deployment and monitoring methods can no longer support large-scale operations. This creates two urgent challenges: how to accelerate the transition of models from experimental labs to production environments, and how to ensure the ongoing reliability and business value of these deployed models.

MLOps addresses these challenges by establishing standardized workflows and automated toolchains. These streamline the entire lifecycle, including data preparation, model training, evaluation, testing, deployment, and monitoring. By leveraging the MLOps capabilities of Huawei Cloud ModelArts, enterprises can achieve agile iteration in model development and intelligent management of production and operations. This approach effectively transforms AI capabilities into tangible service productivity.

Advantages
    • End-to-end automation: MLOps establishes a fully automated pipeline that spans data preprocessing, model training, evaluation, and deployment. This automation minimizes manual intervention and significantly accelerates the transition from development to rollout.

    • Version traceability: By unifying the management of datasets, code, model artifacts, and configurations, MLOps ensures full reproducibility. This traceability simplifies debugging, supports audit compliance, and enables reliable rollback capabilities.

    • Continuous monitoring and adaptive optimization: MLOps systems continuously track online model performance and data distribution shifts. Upon detecting model drift or degradation in accuracy, the system can automatically trigger retraining or incremental updates to maintain service quality.

    • Environment consistency: Leveraging containerization and standardized runtime environments, MLOps eliminates discrepancies between development and production settings. This consistency prevents common deployment failures, such as the "it works on my machine" issue, and improves overall deployment reliability.

    • Elastic resource scheduling: Computing resources are dynamically allocated based on real-time training demands and inference loads. By efficiently utilizing heterogeneous hardware, including GPUs and NPUs, MLOps optimizes hardware utilization and reduces operational costs.

Use Cases
    • FinTech risk control teams face rapidly evolving fraud patterns and risk profiles, making frequent model updates essential for effective credit scoring and fraud detection. Traditional manual deployment often leads to latency in updating models, compromising risk mitigation. MLOps enables automated pipelines that facilitate scheduled daily retraining and grayscale releases to production. Real-time monitoring of key metrics such as recall and precision ensures that risk control policies remain effective, thereby minimizing bad debt losses and enhancing financial security.

    • E-commerce platform algorithm engineers are responsible for maintaining models of the offering recommendation and user profile systems. Facing massive user behavior data and changing market trends, models need to be quickly iterated and optimized. With the continuous integration and continuous deployment capabilities of MLOps, engineers can develop multiple recommendation policy models in parallel. Integrated A/B testing frameworks automatically evaluate model performance, selecting the optimal version for rollout. This agility significantly improves click-through rates and user retention.

    • Manufacturing quality inspection departments deploy visual defect detection models on production lines. However, these models often struggle with generalization across different production lines and product types due to environmental variations. MLOps facilitates multi-model management, enabling quality teams to maintain and monitor dozens of specialized models simultaneously. Centralized tracking of false positive and false negative rates allows for proactive maintenance. When production line processes change, the system can automatically trigger incremental training and version updates, ensuring the detection system remains accurate and effective.

    • Smart campus O&M teams are responsible for the stable running of multiple AI subsystems, including security monitoring and energy management, which face challenges such as model degradation, hardware failures, and data anomalies. MLOps provides a unified operations dashboard that offers centralized visibility into model health, inference latency, and resource consumption. They can define threshold-based alerting rules and automate contingency plans for exception handling. This approach reduces manual inspection workload and enhances the overall intelligence and reliability of campus operations.

    • Medical imaging departments utilize deep learning for diagnostic assistance in areas such as pulmonary nodule detection and retinopathy screening. These models require periodic optimization based on new clinical data and must comply with strict regulatory standards for medical device certification. MLOps supports rigorous version management and audit trails, documenting training data sources, hyperparameter configurations, and validation results. This comprehensive documentation accelerates compliance approvals for model iterations and ensures the safety, efficacy, and regulatory adherence of clinical decision-support tools.

History

The roots of MLOps lie in the DevOps movement within software engineering. Around 2010, as internet applications grew more complex, a collaboration gap emerged between development and operations teams. DevOps addressed this by bridging the conflict between rapid software iteration and stable deployment through continuous integration and continuous delivery (CI/CD).

By 2015, AI and machine learning technologies had gained widespread adoption across industries. However, ML projects possess unique characteristics: they rely on vast datasets, involve stochastic training processes, suffer from significant differences between offline and online environments, and experience model degradation over time. Traditional software engineering methods could not fully accommodate these traits. Consequently, most ML models remained in experimental stages, unable to transition into production systems capable of continuous operation. This limitation prompted the industry to explore extending the DevOps paradigm to the machine learning domain.

In 2017, Google published the paper "Hidden Technical Debt in Machine Learning Systems", which systematically analyzed technical debt in ML systems and outlined specific architectural requirements. That same year, the term MLOps was officially coined, marking a shift from theoretical discussion to practical application. Early MLOps efforts primarily focused on model version management and establishing basic CI/CD pipelines.

From 2019 to 2020, as open-source projects like TensorFlow Extended (TFX) and Kubeflow matured, the MLOps toolchain gradually improved, covering more phases such as data processing, feature storage, experiment tracking, model registration, and service deployment. Major cloud vendors launched managed MLOps services, lowering the barrier for enterprises to adopt MLOps. The core feature of MLOps during this period was tool integration and platformization.

After 2021, MLOps entered the standardization and intelligence phase. The LF AI&Data under the Linux Foundation released the MLOps Maturity Model, and the industry reached a consensus on the classification of MLOps capabilities. At the same time, the rise of LLMs and generative AI has posed new challenges to MLOps: a sharp increase in model scale, high inference costs, and higher requirements for content security and controllability. As a sub-domain of MLOps, Large Language Model Operations (LLMOps) focuses on fine-tuning, evaluation, and governance of LLMs.

Currently, MLOps has become a standard configuration for enterprise AI infrastructure. Huawei Cloud ModelArts deeply integrates MLOps best practices and provides a visualized development environment and automated O&M capabilities throughout the lifecycle. It also collaborates with the Ascend AI computing ecosystem to provide efficient, stable, and secure AI engineering support for customers across various industries.

Key Components

The MLOps system consists of five core modules: data and feature management, model development and training, model registration and version management, deployment and servitization, and monitoring and governance, forming closed-loop management from data to services.

  • The data and feature management module is responsible for collecting, cleaning, labeling, and performing feature engineering on training data. This module provides data version control, lineage tracking, and quality verification mechanisms, and supports feature definition, storage, and reuse. As a core component, the Feature Store centrally manages the feature data required for offline training and online inference, ensuring consistency between training and prediction.

  • The model development and training module provides model design, experiment management, and distributed training capabilities. Developers can develop models using notebooks, visualized orchestration, or code submission. The system automatically records the code version, hyperparameter configuration, evaluation metrics, and output model for each experiment. This module supports single-node and distributed training job scheduling, fully utilizing cluster computing resources to accelerate model convergence.

  • The model registration and version management module serves as the central repository for model assets, enabling unified registration, classification, and versioning of trained models. Each model version is associated with complete metadata information, including the training dataset, performance metrics, application scenarios, and compliance approval status. This module supports the model review and approval process, ensuring that only verified models are eligible for the production deployment candidate list.

  • The deployment and servitization module publishes verified models as real-time API services or batch inference jobs. This module supports multiple deployment modes: real-time inference services are suitable for scenarios requiring low latency, batch inference is ideal for large-scale offline data processing, and edge deployment meets localized computing needs. By using blue-green deployment and canary release strategies, this module ensures smooth model updates and zero service interruption.

  • The monitoring and governance module performs comprehensive health checks on online models, covering service metrics such as inference latency, throughput, and error rate, as well as model quality metrics like prediction distribution, feature shift, and label drift. When an anomaly is detected, the system automatically sends an alert and triggers a rollback or retraining based on predefined policies. This module also provides model usage audit, access control, and cost analysis functions to meet enterprise governance and compliance requirements.

The five modules are closely connected through a pipeline engine. The data module outputs a feature set, which is then input into the training module to generate candidate models. After version management and approval by the registration module, the deployment module releases the models as production services. The monitoring module continuously collects online feedback data and feeds it back to the data module, forming a closed loop. This design enables automated cycling and continuous optimization throughout the model lifecycle.

Figure1 MLOps components

Based on this modular architecture, MLOps enables the transition of machine learning from experimental exploration to industrial-scale production, providing enterprises with a scalable, reliable, and continuously optimized AI capability delivery system.

How MLOps Works

The operation of the MLOps system begins when a user defines a machine learning task and triggers the pipeline execution. Users can initiate requests through console configuration, API calls, or code commits. The input objects include the original dataset, training script, hyperparameter configuration, and deployment policy. After receiving the request, the system automatically advances tasks in each phase according to the preset process.

Step 1: Prepare data and perform feature engineering.

The pipeline pulls the raw data of the specified version from the data storage and performs preprocessing operations such as data cleansing, missing value imputation, and outlier handling. Then, the system generates a training feature set based on predefined feature transformation rules and writes the processed features to the feature storage. This step ensures data quality and feature consistency, providing reliable input for subsequent training.

Step 2: Train the model and record experiments.

The system invokes a distributed training cluster to load feature data and performs model training based on the specified algorithm and hyperparameter configurations. During training, the loss curve, evaluation metrics, and resource usage are recorded in real time. After the training is complete, the generated model artifacts (including the network structure, weight parameters, and inference graph) are packaged and stored. The experiment metadata is synchronized to the experiment management platform for comparison and analysis.

Step 3: Evaluate the model and register the version.

The pipeline automatically evaluates the model performance on the validation and test sets and compares the metrics of the baseline model and the historical best model. If the new model meets the preset admission criteria (for example, the accuracy improvement exceeds the threshold), it is registered in the model repository and assigned a version number. In scenarios that require manual review, the system triggers the approval workflow and waits for domain experts to confirm.

Step 4: Perform continuous integration and deployment.

The approved model enters the deployment phase. The system first encapsulates the model into a standard container image and injects runtime dependencies and configuration information. Then, inference service endpoints are created based on the deployment policy. Both the old and new versions are run simultaneously, and the migration is completed through traffic switching. The traffic proportion of the new version is gradually increased, and the stability is observed. Batch inference jobs are scheduled and executed offline as planned.

Step 5: Perform online monitoring and collect metrics.

The deployed model receives real service requests. The monitoring system collects service-level metrics such as inference latency, QPS, and error rate in real time. It also samples and saves input features and output results for subsequent analysis. The data drift detection module periodically compares the statistical distribution of newly released data with the training data baseline to identify potential distribution shift risks.

Step 6: Perform alarm decision-making and automatic triggering.

When a monitored metric exceeds the preset threshold or significant data drift is detected, the alarm module sends a notification to the O&M team. Based on the preset automation policies, the system can make autonomous decisions: slight performance fluctuations trigger log recording and manual review; significant accuracy drops trigger the incremental training pipeline; severe faults trigger model rollback to the previous stable version. A new training job will be initiated with the latest data to restart the preceding process, forming a continuous optimization loop.

Figure1 MLOps workflow

The preceding steps form a complete closed-loop operation of MLOps: data-driven model training, model deployment after strict verification, and online performance feedback to data accumulation and the next round of optimization. Through automated pipelines and intelligent decision-making mechanisms, MLOps ensures that models continuously adapt to business requirements throughout their lifecycle, providing users with stable and reliable AI service capabilities.

Types of MLOps

Based on the implementation depth and application scenarios, MLOps can be classified from two dimensions: maturity level and service deployment mode.

 

Table1 By maturity

Type

Core Feature

Typical Application Scenario

Targeted Advantages

MLOps 1.0 - Manual management

Model development and deployment currently rely on manual processes, lacking an automated toolchain. Consequently, version management depends on manual records.

Small-scale AI projects, prototype verification, and research topics.

High flexibility, no need for additional tools, and suitable for exploratory work.

MLOps 2.0 - Pipeline automation

CI/CD pipelines are built to automate training and deployment, with basic version management and experiment tracking capabilities.

Medium-sized enterprises that require concurrent development of multiple models and regular iterative updates.

Significantly improves efficiency, reduces human errors, and supports team collaboration.

MLOps 3.0 - Continuous intelligence and governance

Comprehensive automation covers data versions, feature management, model monitoring, automatic retraining, and built-in compliance audit and governance mechanisms.

Large financial institutions, medical imaging, autonomous driving, and other scenarios with high compliance and reliability requirements.

Achieves true continuous learning and adaptive optimization, meeting strict regulatory requirements.

Table1 By service deployment mode

Type

Core Feature

Application Scenario

Targeted Advantages

Cloud-hosted MLOps

The full-process toolchain is provided by cloud service providers, eliminating the need for users to build their own infrastructure and allowing for pay-per-use.

Startups, rapid trial-and-error projects, and resource-limited teams.

Reduces initial investment, provides out-of-the-box functionality, and offers elastic compute power and O&M support from the cloud platform.

On-premises MLOps

The MLOps platform is deployed in an enterprise's own data center or private cloud, ensuring complete data isolation.

Scenarios where government agencies, financial institutions, and confidential industries have strict requirements on data security.

Data remains within the domain, meeting compliance and audit requirements, and offers high customization.

Hybrid cloud-based MLOps

Training utilizes elastic compute on the cloud, while inference is executed at the edge or locally. Both are managed through a unified platform.

Scenarios that require local low-latency response, such as smart manufacturing, smart campus, and retail chain stores.

Balances the advantages of computing resources on the cloud with the response speed at the edge response, effectively combining centralized management with distributed execution.

Figure1  Types of MLOps

This classification empowers organizations to tailor their MLOps strategies to their specific scale, regulatory needs, and business contexts, enabling a structured evolution of their AI engineering capabilities.

 

How Huawei Cloud Supports Your MLOps Needs

Huawei Cloud has developed a comprehensive AI development and O&M product line based on the core concepts of MLOps, helping enterprises efficiently build and operate machine learning systems.

  • ModelArts, a one-stop AI development platform, natively supports the entire MLOps process and provides an integrated workbench for data annotation, visual modeling, code development, distributed training, model management, and real-time service deployment. The platform has a built-in workflow engine that allows users to build automated pipelines through drag-and-drop operations, achieving end-to-end automation from data import to model publishing. ModelArts supports multi-tenant collaboration and fine-grained permission control, meeting the collaboration requirements of enterprise-level teams. 

    You can create your first MLOps project by referring to ModelArts Service Overview to experience the automated model development process.

  • SoftWare Repository for Container (SWR) provides efficient containerization support for MLOps. Both the training environment and inference service can be encapsulated into standard Docker images. SWR provides image scanning, version management, and global accelerated distribution capabilities to ensure consistent delivery of model services across different environments. Working with Cloud Container Engine (CCE), SWR enables auto scaling and grayscale release of inference services. For details, see Getting Started with SWR.

  • Log Tank Service (LTS) and Application Operation Management (AOM) form the monitoring and alarm base of MLOps. LTS centrally collects training logs and inference access logs, and supports keyword search and real-time monitoring dashboards. AOM provides application topology, performance metrics, and tracing capabilities, helping O&M teams quickly locate bottlenecks and exceptions. Working with Cloud Eye, you can set multi-dimensional threshold alarm rules and notify relevant personnel by SMS, email, or webhook. For details, see Accessing LTS.

Related Products

Related Products