Updated on 2026-08-28 GMT+08:00

Text Generation SFT

Centered on Q&A-style annotated text datasets, this format contains clear instructions and standard response content. This type is deeply adapted for supervised instruction fine-tuning of LLMs, meeting model training requirements for dialogue generation, instruction alignment, and custom scenario adaptations in specific industries.

Constraints

  • Import from OBS: The size of a single file or compressed package cannot exceed 20 GB. If multiple files are imported, the total file size cannot exceed 20 GB.
  • Local import: The size of a single file cannot exceed 1 GB, and the number of files cannot exceed 20.
  • File encoding: All JSONL format files must use UTF-8 encoding.

Format Requirement

File formats: JSONL

Format description: JSONL is a structured, lightweight text file format. Its core rule is that every line in the file must be an independent, valid, and complete JSON object.

Dataset logical specifications: To accommodate the conventions of different open-source communities, this data type supports three mainstream logical specifications: standard format, Alpaca format, and ShareGPT format.

Standard Format

This specification follows industry-standard conventions and records conversations using a message sequence, offering exceptionally high versatility and compatibility.

Data format example

To make the field hierarchy clearer, the single-line JSONL data has been expanded.

{
  "meta": {
    "conversation_id": "conversation_123",
    "user_id": "user_456",
    "timestamp": "2026-05-29T10:00:00Z",
    "domain": "math_tutoring"
  },
  "messages": [
    {
      "role": "user",
      "content": "Please help me solve the equation 2x + 5 = 13."
    },
    {
      "role": "assistant",
      "content": "<think>This is a linear equation in one variable. I need to move the constant term to the right side and then divide by the coefficient.</think>\nSure, let's solve this equation: 2x + 5 = 13. Subtracting 5 from both sides gives 2x = 8, so x = 4."
    },
    {
      "role": "user",
      "content": "What about 3(x - 2) = 9?"
    },
    {
      "role": "assistant",
      "content": "<think>This equation contains parentheses. I can either expand them first or divide both sides by 3 directly.</think>\nYou can solve it like this: dividing both sides by 3 gives x - 2 = 3, so x = 5."
    },
    {
      "role": "user",
      "content": "Thanks, I understand now."
    },
    {
      "role": "assistant",
      "content": "You're welcome! Feel free to ask if you have any more questions."
    }
  ]
}

Field specifications

The table below shows an example of each entry (line) in the JSONL file.

Table 1 Fields

Field

Type

Mandatory

Description

messages

List

Yes

The dialogue history list, containing multi-turn role interactions in chronological order.

messages.role

String

Yes

The role identity of the message sender. Only supports one of the following: system (persona/system prompt), user (user input), or assistant (model response).

messages.content

String

Yes

The specific conversation text content. If it includes the model's thinking process, <think>...</think> tags can be used to wrap it.

meta

-

No

Extended metadata field.

Alpaca Format

This specification originates from the Stanford Alpaca community project. It breaks down single interactions or multi-turn dialogues with history into clear instructions (instruction), supplementary context (input), and standard output (output). Its flat structure makes it easy to read and suitable for rapid fine-tuning.

Data format example

To make the field hierarchy clearer, the single-line JSONL data has been expanded.

{
  "system": "You are an AI teaching assistant proficient in elementary mathematics.",
"instruction": "Please help me solve the equation 2x + 5 = 13.",
  "input": "",
  "output": "<think>This is a linear equation in one variable. I need to move the constant term to the right side and then divide by the coefficient.</think>\nSure, let's solve this equation: 2x + 5 = 13. Subtracting 5 from both sides gives 2x = 8, so x = 4."
  "history": [
    [
      "Hello, can you help me tutor math?",
      "Of course! Feel free to ask any math questions you have."
    ]
  ]
}

Field specifications

The table below shows an example of each entry (line) in the JSONL file.

Table 2 Fields

Field

Type

Mandatory

Description

system

String

No

The system prompt used to explicitly define the persona/role assumed by the model during this conversation.

instruction

String

Yes

The specific task instruction or core question issued by the user.

input

String

Yes

Supplementary background information or input context required for the task (if none, an empty string "" can be passed).

output

String

Yes

The standard expected response from the LLM for the given instruction. It may include a chain-of-thought reasoning process wrapped in <think>...</think> tags.

history

List

No

Historical conversation records, structured as a two-dimensional list [[q1, a1], [q2, a2], ...]. Each element is a list with a fixed length of 2, representing the user question and model response in a historical turn, respectively.

ShareGPT Format

This specification is widely used across various open-source frameworks. It linearly reproduces the multi-turn alternating interaction process between humans and AI through a centralized conversation list (conversations), featuring a compact structure and exceptional extensibility.

Data format example

To make the field hierarchy clearer, the single-line JSONL data has been expanded.

{
  "system_prompt": "Role: Mathematics Tutoring Expert",
  "conversations": [
    {
      "from": "human",
      "value": "Please help me solve the equation 2x + 5 = 13."
    },
    {
      "from": "gpt",
      "value": "<think>Transpose terms, simplify, and solve.</think>\nSolving the equation gives 2x = 8, so x = 4."
    }
  ]
}

Field specifications

The table below shows an example of each entry (line) in the JSONL file.

Field

Type

Mandatory

Description

conversations

List

Yes

An ordered list of multi-turn dialogue content.

conversations.from

String

Yes

Identifies the speaker of the current utterance.

conversations.value

String

Yes

The text-based dialogue content for the current turn.