{"_id":"6662064bb177c2027457665e","id":"livebench/data_analysis","author":"livebench","sha":"31b9661ff678df9958e2f7fa228427f4c858c1a1","lastModified":"2025-04-07T20:34:15.000Z","private":false,"gated":false,"disabled":false,"tags":["size_categories:n<1K","format:parquet","modality:text","library:datasets","library:pandas","library:mlcroissant","library:polars","arxiv:2406.19314","region:us"],"description":"\n\t\n\t\t\n\t\n\t\n\t\tDataset Card for \"livebench/data_analysis\"\n\t\n\nLiveBench is a benchmark for LLMs designed with test set contamination and objective evaluation in mind. It has the following properties:\n\nLiveBench is designed to limit potential contamination by releasing new questions monthly, as well as having questions based on recently-released datasets, arXiv papers, news articles, and IMDb movie synopses.\nEach question has verifiable, objective ground-truth answers, allowing hard questions to be… See the full description on the dataset page: https://huggingface.co/datasets/livebench/data_analysis.","downloads":5384,"likes":7,"cardData":{"dataset_info":{"features":[{"name":"question_id","dtype":"string"},{"name":"category","dtype":"string"},{"name":"turns","sequence":"string"},{"name":"ground_truth","dtype":"string"},{"name":"task","dtype":"string"},{"name":"livebench_release_date","dtype":"timestamp[s]"},{"name":"livebench_removal_date","dtype":"string"}],"splits":[{"name":"test","num_bytes":305848,"num_examples":150}],"download_size":144796,"dataset_size":305848},"configs":[{"config_name":"default","data_files":[{"split":"test","path":"data/test-*"}]}],"arxiv":2406.19314},"siblings":[{"rfilename":".gitattributes"},{"rfilename":"README.md"},{"rfilename":"data/test-00000-of-00001.parquet"}],"createdAt":"2024-06-06T18:56:11.000Z","usedStorage":1327264}