Overview
Core Features
Smart refining is a key part of ModelArts data engineering, designed to address the dual challenges of data quality and quantity in foundation model training. It transcends the boundaries of traditional data processing by seamlessly integrating rule-based data processing (cleaning, filtering, deduplication, and more) with LLM-based data synthesis (rewriting, expansion, polishing, and more).
Through visualized operator orchestration, you can drag and drop multiple processing and synthesis operators to build an automated pipeline, much like assembling building blocks. The system follows your predefined logic to filter and optimize massive raw datasets layer by layer, ultimately outputting high-quality datasets that meet rigorous training requirements.
Functional Architecture
Smart refining takes text, image, and video datasets as input. It constructs smart refining tasks by orchestrating various data processing and synthesis operators to produce refined datasets. For details, see Figure 1.
Benefits
- Unified workflow: Orchestrates data processing and synthesis in a single pipeline, eliminating the need to switch between modules and reducing intermediate data transfers.
- Enhanced data quality: Ensures high-quality input for the synthesis stage through rigorous, multi-level filtering using processing operators.
- Flexible orchestration: Supports the free combination of dozens of operators to satisfy diverse scenarios, from simple cleaning to complex data augmentation.
- Efficient scale expansion: Enables high-efficiency training data expansion by performing synthetic rewriting on top of cleaned, high-quality data.
- Streamlined operation: Offers a visualized, "what you see is what you get" orchestration experience, removing the need for manual scripting.
- Workflow reproducibility: Supports saving and reusing refining templates to ensure consistency across different data processing tasks.
Scenarios
Smart refining is used in these typical scenarios, each with recommended operator combinations. Choose a scenario based on your needs.
- Choose Scenario 1: Converting the Dataset Format if you only need to convert the format of text datasets.
- Choose Scenario 2: Cleaning and Improving Quality of Raw Corpus if the data quality is poor and needs to be cleaned.
- Choose Scenario 3: Expanding and Augmenting Training Data if the data volume is insufficient and needs to be expanded.
- Choose Scenario 4: Preparing Data for SFT to prepare SFT data.
- Choose Scenario 5: Processing Multimodal Data to process image or video data.
- Choose Scenario 6: Ensuring Data Compliance and Security if you need to meet data compliance requirements.
Scenario 1: Converting the Dataset Format
Description
ModelArts supports multiple dataset formats. You need to convert data from one format to another without additional data processing.
Recommended operator orchestration sequence
Start node → end node
Expected results
The input data is converted into output data in different formats (standard, Alpaca, or ShareGPT).
Scenario 2: Cleaning and Improving Quality of Raw Corpus
Description
Raw data from the internet, internal systems, or third parties often has a lot of noise. So, it needs to be systematically cleaned for model training. Common corpus issues are listed in the table below.
| Data Issue | Form | Impact on the Model |
|---|---|---|
| Duplicate information | A large amount of identical or similar content exists in the data. | The trained model is overfitting. |
| Garbled data | Encoding errors and abnormal characters exist in the data. | The semantic understanding of the model is polluted. |
| Sensitive and non-compliant information | Political, pornographic, or violent content exists in the data. | The model output has compliance risks. |
| Poor data quality | Sentences are not coherent, the logic is disordered, and sentences are incomplete. | The model generation quality is reduced. |
| Invalid data length | The data is too short to be meaningful or too long to be redundant. | The training efficiency is low. |
| Mixed data, not classified | Data from various domains is mixed and not classified by domain. | Unclassified domain data affects training efficiency. |
Recommended operator orchestration sequence
Raw corpus → [Symbol standardization] → [Deduplication operator] → [Sensitive word filtering] → [Text length filtering] → [Incomplete sentence removal at paragraph ends] → [Pornographic text detection] → [Political text detection] → [Insult text detection operator] → [Pre-trained text classification] → Cleaned data
Expected results
- The data repetition rate is reduced by more than 90%.
- Low-quality samples are effectively removed.
- 100% of sensitive and non-compliant content is filtered out.
- The output data can be directly used for training or proceed to the next synthesis step.
- The output data can be classified by domain.
Scenario 3: Expanding and Augmenting Training Data
The cost of obtaining high-quality labeled data is high, and the existing data volume is insufficient to train a model with good performance.
Applicable scenarios
- Data is scarce in vertical domains.
- The labeling cost is too high.
- The data scale needs to be quickly expanded.
- Data diversity is insufficient.
Recommended operator orchestration sequence
Raw data → Data cleaning → Data generation → Expanded data
| Operator | Function | Configuration Suggestion |
|---|---|---|
| Data cleaning | Ensures seed data quality. | Ensure strict screening criteria. |
| Data generation | Generates diverse expressions. | Select an appropriate rewriting strategy to generate diverse data. |
Expected results
- The semantic consistency is maintained.
- The expression diversity is improved.
Scenario 4: Preparing Data for SFT
Prepare high-quality datasets for SFT of foundation models.
Applicable scenarios
- Fine-tuning of general assistant models
- Customization of industry-specific models
- Dialogue capability optimization
- Task-oriented model training
Recommended operator orchestration process
Raw instruction data → Data cleaning → Text generation (optional) → Dataset generation
| Input Format | Processing Method | Output Format |
|---|---|---|
| Unstructured text | Format conversion operator | Alpaca/ShareGPT |
| Existing Alpaca | Quality filtering + rewriting | Optimized Alpaca |
| Existing ShareGPT | Quality filtering + rewriting | Optimized ShareGPT |
Key quality control points
- Instruction clarity check
- Answer accuracy verification
- Format consistency assurance
Scenario 5: Processing Multimodal Data
Process datasets that contain multiple modalities, such as images and videos.
Applicable scenarios
Video understanding data sorting
Recommended operator orchestration process (using images as an example)
Image dataset → Image deduplication → Image extraction → Image metadata filtering → Image detection → Processed data
| Modality | Key Point | Notes |
|---|---|---|
| Image | Size, format, and quality | Unified resolution |
| Video | Frame rate, resolution, and segment | Unified video encoding |
Important constraints
A data processing task needs to be created for each modality separately.
Scenario 6: Ensuring Data Compliance and Security
Ensure that the training data complies with regulatory requirements and enterprise security policies.
Applicable scenarios
- Personal information protection (GDPR/Personal Information Protection Law)
- Sensitive data filtering
Recommended operator orchestration process
Raw data → Sensitive word filtering → Compliance data
Operators to use
| Operator | Function | Compliance Requirements |
|---|---|---|
| Sensitive word filtering | Filters sensitive content in personal information. | Personal privacy requirements must be met. |
What is your overall rating for this page?
Thank you very much for your feedback. We will continue working to improve the documentation.See the reply and handling status in My Cloud VOC.
For any further questions, feel free to contact us through the chatbot.
Chatbot
