Updated on 2026-08-28 GMT+08:00

Preset Data

ModelArts provides high-quality preset datasets for you. These datasets follow open-source rules and work with popular training frameworks. They help track dataset versions and reproduce experiments. Choose a suitable dataset for your needs and use it right away in the platform.

Scenarios

Typical scenarios of Preset Data:

  • Use preset datasets along with your own data to complete smart refining to improve and create high-quality datasets for later tasks.
  • Use a preset dataset for LLM pre-training and fine-tuning to enhance foundational capabilities, while leveraging human preference data to optimize response quality.
  • Combine image, video, and audio data to build cross-modal capabilities and multimodal models.
  • Use a dataset as a standard test set to evaluate model performance and establish baseline assessments of model capabilities.

Viewing Preset Data

  1. Log in to the ModelArts console.
  2. In the navigation pane, choose Asset Management > Data. Click the Preset Data tab. The preset datasets are displayed in cards.

    The preset data card displays the data name, type, description, update time, and number of samples.

  3. Click a preset dataset card to view its details. The details include Basic Info and Data Preview.
    • Basic Info: displays the dataset name, type, number of samples, dataset size, and description of the preset dataset.
      Figure 1 Basic information about preset data

    • Data Preview allows you to display some typical samples of structured data (text and tables), view the samples on multiple pages, and view the original data structure. Unstructured data (images/audio) can be previewed in thumbnail mode.
      Figure 2 Preset data preview

Preset Datasets

ModelArts offers preset text and image datasets. For details, see Table 1. Choose a dataset that fits your scenario.

Table 1 List of preset datasets

Name

Data Modality

Dataset Type

Dataset Format

Dataset Overview

Size

Samples

Language

Link

GPT-4-LLM

Text

Text generation SFT

Alpaca format

Alpaca-CoT is a large, high-quality dataset for instruction fine-tuning that includes various task types.

32.6 MB

48,818

Chinese

GPT-4-LLM

gsm8k

Other

Custom

-

GSM8K (Grade School Math 8K) is a dataset containing 8,500 high-quality, linguistically diverse grade school math word problems. Designed to support question-answering tasks for basic math problems, the dataset requires multi-step reasoning.

1.2 MB

2

English

gsm8k

ArcherCodeR

Other

Custom

-

A code dataset constructed based on multiple open-source datasets, including deepcoder-preview-datasets, deepmind, code_contests, open-r1, and codeforces, used for reinforcement learning training.

6.9 GB

2

English

ArcherCodeR

hotpot_qa

Other

Custom

-

HotpotQA is a question-answering dataset collected from English Wikipedia, containing approximately 113K crowdsourced questions that require the introductory paragraphs of two Wikipedia articles to answer.

8.4 MB

2

English

hotpot_qa

Geometry3K

Other

Custom

-

A new large-scale geometry problem-solving dataset containing 3,002 multiple-choice geometry problems with dense formal language annotations in both diagrammatic and textual formats: 27,213 annotated diagrammatic logical forms (literals) and 6,293 annotated textual logical forms (literals).

51.1 MB

2

English

Geometry3K

code_alpaca

Text

Text generation SFT

Alpaca format

This dataset was released by CodeAlpaca and contains code generation tasks with 20,022 samples.

6.7 MB

20,022

English

code_alpaca

alpaca_gpt4_data

Text

Text generation SFT

Alpaca format

This dataset was released by Instruction-Tuning-with-GPT-4. It contains 52,000 English instruction-following samples generated by GPT-4 using Alpaca prompts, and is used to fine-tune LLMs.

40.4 MB

52,002

English

alpaca_gpt4_data

alpaca_data

Text

Text generation SFT

Alpaca format

Stanford Alpaca released this dataset, which has 52,000 English instruction samples created using self-supervised methods.

20.0 MB

52,002

English

alpaca_data

ai-expert-alpaca

Text

Text generation SFT

Alpaca format

This dataset contains high-quality Q&A pairs for SFT of LLMs, focusing on three core AI technology domains: LLMs, retrieval-augmented generation (RAG), and agent systems. The dataset covers these advanced AI topics in both English and Chinese.

8.2 MB

11,235

-

ai-expert-alpaca