Datasets: NeurIPS LLM Challenge 2023
Datasets that were under consideration for usage in my submission to the 2023 NeurIPS Large Language Model Efficiency Challenge.
Viewer • Updated • 6.09k • 32Note Ultimately used in my full eval submission, with exclusion of dolly_hhrlhf. Included only in Mistral-7B-sft-v1.
databricks/databricks-dolly-15k
Viewer • Updated • 19.4k • 643Note Used both for Mistral-7B-sft-v0 and Mistral-7B-sft-v1 in my submissions.
hendrycks/competition_math
Updated • 8.98k • 88Note This turned out to be one of the holdout tasks. mosaicml/instruct-v3 includes training problems from competition_math. This is a contributor to the high scores on GSM8k and MATH benchmarks.
kaist-ai/CoT-Collection
Viewer • Updated • 373 • 85Note Looked promising, but did not have time to explore.
tasksource/icl-symbol-tuning-instruct
Viewer • Updated • 6 • 16Note Considered for improving ICL. Did not have time to explore.
cais/mmlu
Viewer • Updated • 1.5M • 224Note Decided against training on MMLU data.
GAIR/lima
Viewer • Updated • 5.41k • 377Note Avoided due to CC BY-NC-SA license, though it would have been allowed for the competition. Likely would have been a good resource otherwise.
grammarly/coedit
Viewer • Updated • 2.55k • 45Note The plan here would be to target robustness metrics by finetuning an expert model to correct perturbations and/or clarify the input. This could have paraphrasing or other text revision tasks if they appeared in the hidden eval. Did not have time to fully explore.
leslyarun/c4_200m_gec_train100k_test25k
Viewer • Updated • 62 • 5Note Similar use case as coedit.
wanyu/IteraTeR_human_sent
Viewer • Updated • 27Note Similar use case as coedit.
social_i_qa
Viewer • Updated • 14.5k • 7Note Now knowing that the holdout tasks had ethics questions, I wish I had used this.
lighteval/siqa
Viewer • Updated • 88.9k • 3Note Same as social_i_qa
tau/commonsense_qa
Viewer • Updated • 7.05k • 51Note Now knowing that the holdout tasks had ethics questions, I wish I had used this.
euirim/goodwiki
Viewer • Updated • 161 • 47Note Could have been useful for RAG.
multi_news
Viewer • Updated • 2.27k • 57Note The thought was this could help with CNN/DM summarization, but some quality and license concerns combined with acceptable performance without it led to its exclusion.
math_qa
Viewer • Updated • 19.6k • 68allenai/ropes
Viewer • Updated • 84 • 40allenai/openbookqa
Viewer • Updated • 1.57k • 60allenai/ai2_arc
Viewer • Updated • 704k • 83riddle_sense
Viewer • Updated • 213 • 21allenai/qasc
Viewer • Updated • 157 • 9nyu-mll/blimp
Viewer • Updated • 1.24k • 31google/boolq
Viewer • Updated • 5.48k • 53corypaik/prost
Viewer • Updated • 363 • 1allenai/sciq
Viewer • Updated • 3.81k • 75facebook/belebele
Viewer • Updated • 20.8k • 76derek-thomas/ScienceQA
Viewer • Updated • 1.32k • 107openlifescienceai/medmcqa
Viewer • Updated • 2.23k • 94embedding-data/QQP_triplets
Viewer • Updated • 439 • 4VMware/open-instruct
Viewer • Updated • 149 • 39