Updated on 2026-08-28 GMT+08:00

Base Text

Centered on multi-format raw documents, this feature supports deep extraction and parsing of document content. Once uploaded, the base text data can be transformed into pre-training datasets or evaluation sets using the platform's built-in data refinement feature.

Constraints

  • Import from OBS: The size of a single file or compressed package cannot exceed 20 GB. If multiple files are imported, the total file size cannot exceed 20 GB.
  • Local import: The size of a single file cannot exceed 1 GB, and the number of files cannot exceed 20.

Format Requirement

File formats: DOCX, PDF

Format description: This type directly handles raw document data, which typically cannot be read directly by models for training. After connection, it must go through a data refinement task to be converted into pre-training text-format datasets before it can be used by downstream tasks.

Document guidelines: To ensure the extraction accuracy and conversion quality of downstream data refinement, the uploaded raw documents should have a clear paragraph structure and recognizable text hierarchies. Avoid using entirely image-based or scanned PDF files without OCR as much as possible. Complex layouts inside documents, such as tables and formulas, will be converted into flattened text descriptions during the refinement stage; note that high-density charts and tables may cause semantic discontinuities.

Figure 1 Example of base text