Base Text
Centered on multi-format raw documents, this feature supports deep extraction and parsing of document content. Once uploaded, the base text data can be transformed into pre-training datasets or evaluation sets using the platform's built-in data refinement feature.
Constraints
- Import from OBS: The size of a single file or compressed package cannot exceed 20 GB. If multiple files are imported, the total file size cannot exceed 20 GB.
- Local import: The size of a single file cannot exceed 1 GB, and the number of files cannot exceed 20.
Format Requirement
File formats: DOCX, PDF
Format description: This type directly handles raw document data, which typically cannot be read directly by models for training. After connection, it must go through a data refinement task to be converted into pre-training text-format datasets before it can be used by downstream tasks.
Document guidelines: To ensure the extraction accuracy and conversion quality of downstream data refinement, the uploaded raw documents should have a clear paragraph structure and recognizable text hierarchies. Avoid using entirely image-based or scanned PDF files without OCR as much as possible. Complex layouts inside documents, such as tables and formulas, will be converted into flattened text descriptions during the refinement stage; note that high-density charts and tables may cause semantic discontinuities.
What is your overall rating for this page?
Thank you very much for your feedback. We will continue working to improve the documentation.See the reply and handling status in My Cloud VOC.
For any further questions, feel free to contact us through the chatbot.
Chatbot