يُعرض النص الكامل باللغة الإنجليزية؛ وقد تمت ترجمة العنوان والملخص.
ما هو
Datasets is an open-source library from Hugging Face, released in 2020, for loading, processing and sharing datasets. It uses Apache Arrow as its underlying format and supports memory mapping and streaming, so datasets larger than memory can still be iterated batch by batch. It addresses the difficulty of reusing the download, parsing and preprocessing pipelines that precede training.
لماذا يستحق التذكّر
It made "getting the data" a cacheable, streamable standard step, so preprocessing code travels and reproduces alongside the dataset.
المواصفات الأساسية
- Backing format
- Apache Arrow
- Reading
- Memory mapping and streaming
المفاهيم ذات الصلة
البنية التحتية للتدريب والاستدلال
تحدد الذاكرة حجم النموذج القابل للتدريب، ويحدد الاتصال المدة التي يستغرقها — ونادرًا ما تكون الحوسبة الخام هي عنق الزجاجة
التدريب المسبق والضبط الدقيق
تعلّم اللغة أولاً من نصوص ضخمة بلا وسوم ثم التخصص ببيانات قليلة — أكثر النماذج كفاءةً في البيانات