Der vollständige Artikel liegt auf Englisch vor; Titel und Zusammenfassung sind lokalisiert.
WAS ES IST
Datasets is an open-source library from Hugging Face, released in 2020, for loading, processing and sharing datasets. It uses Apache Arrow as its underlying format and supports memory mapping and streaming, so datasets larger than memory can still be iterated batch by batch. It addresses the difficulty of reusing the download, parsing and preprocessing pipelines that precede training.
Warum es wichtig ist
It made "getting the data" a cacheable, streamable standard step, so preprocessing code travels and reproduces alongside the dataset.
Wichtige Eckdaten
- Backing format
- Apache Arrow
- Reading
- Memory mapping and streaming
Verwandte Konzepte
Infrastruktur für Training und Inferenz
Der Speicher bestimmt, wie groß ein Modell sein darf, die Kommunikation, wie lange das Training dauert
Vor- und Feintraining
Erst Sprache aus riesigem ungelabelten Text lernen, dann mit wenig Daten spezialisieren – das dateneffizienteste Paradigma der modernen KI
Vergleichbare Produkte
Transformers
2018Eine API zum Laden und Trainieren vortrainierter Modelle
Diffusers
2022Einheitliche Umsetzung und Scheduler für Diffusionsmodelle
Pinecone
2019Verwaltete Vektordatenbank für Ähnlichkeitssuche