본문으로 건너뛰기
AI 도감

Datasets

한 줄 코드로 데이터셋 로드와 스트리밍

Hugging Face 도구 오픈 소스

이 페이지의 본문은 영어로 제공됩니다. 제목과 요약은 한국어로 번역되었습니다.

무엇인가

Datasets is an open-source library from Hugging Face, released in 2020, for loading, processing and sharing datasets. It uses Apache Arrow as its underlying format and supports memory mapping and streaming, so datasets larger than memory can still be iterated batch by batch. It addresses the difficulty of reusing the download, parsing and preprocessing pipelines that precede training.

기억할 만한 이유

It made "getting the data" a cacheable, streamable standard step, so preprocessing code travels and reproduces alongside the dataset.

주요 사양

Backing format
Apache Arrow
Reading
Memory mapping and streaming

관련 개념

동종 제품