arxiv:2109.02846

Datasets: A Community Library for Natural Language Processing

Published on Sep 7, 2021

Upvote

Authors:

Quentin Lhoest ,

Albert Villanova del Moral ,

Yacine Jernite ,

Abhishek Thakur ,

Patrick von Platen ,

Suraj Patil ,

Julien Chaumond ,

Lewis Tunstall ,

Joe Davison ,

Mario Šaško ,

Gunjan Chhablani ,

Simon Brandeis ,

Teven Le Scao ,

Victor Sanh ,

Canwen Xu ,

Nicolas Patry ,

Angelina McMillan-Major ,

Philipp Schmid ,

Sylvain Gugger

Abstract

The scale, variety, and quantity of publicly-available NLP datasets has grown rapidly as researchers propose new tasks, larger models, and novel benchmarks. Datasets is a community library for contemporary NLP designed to support this ecosystem. Datasets aims to standardize end-user interfaces, versioning, and documentation, while providing a lightweight front-end that behaves similarly for small datasets as for internet-scale corpora. The design of the library incorporates a distributed, community-driven approach to adding datasets and documenting usage. After a year of development, the library now includes more than 650 unique datasets, has more than 250 contributors, and has helped support a variety of novel cross-dataset research projects and shared tasks. The library is available at https://github.com/huggingface/datasets.

View arXiv page View PDF Add to collection

Community

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.

Tap or paste here to upload images

· Sign up or log in to comment

Upvote

Models citing this paper 0

No model linking this paper

Cite arxiv.org/abs/2109.02846 in a model README.md to link it from this page.

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2109.02846 in a dataset README.md to link it from this page.

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2109.02846 in a Space README.md to link it from this page.