本ページの本文は英語で提供されています。タイトルと導入は日本語化されています。
これは何か
Datasets is an open-source library from Hugging Face, released in 2020, for loading, processing and sharing datasets. It uses Apache Arrow as its underlying format and supports memory mapping and streaming, so datasets larger than memory can still be iterated batch by batch. It addresses the difficulty of reusing the download, parsing and preprocessing pipelines that precede training.
なぜ覚えておく価値があるか
It made "getting the data" a cacheable, streamable standard step, so preprocessing code travels and reproduces alongside the dataset.
主な仕様
- Backing format
- Apache Arrow
- Reading
- Memory mapping and streaming