Перейти к содержимому
Атлас ИИ

Datasets

Загрузка и потоковое чтение наборов данных одной строкой

Hugging Face Инструмент Открытый исходный код

Полный текст статьи представлен на английском; заголовок и аннотация локализованы.

ЧТО ЭТО

Datasets is an open-source library from Hugging Face, released in 2020, for loading, processing and sharing datasets. It uses Apache Arrow as its underlying format and supports memory mapping and streaming, so datasets larger than memory can still be iterated batch by batch. It addresses the difficulty of reusing the download, parsing and preprocessing pipelines that precede training.

Почему стоит запомнить

It made "getting the data" a cacheable, streamable standard step, so preprocessing code travels and reproduces alongside the dataset.

Ключевые характеристики

Backing format
Apache Arrow
Reading
Memory mapping and streaming

Связанные концепции

Аналоги