YAML Metadata Warning:empty or missing yaml metadata in repo card
Check out the documentation for more information.
Paper
This model has been trained as part of our research paper, Does Sensitivity to Pragmatic Norms Emerge with More Parameters, More Data, or More Conversations? accepted for EMNLP Main 2026. For details on the model architecture, training procedure, data preprocessing, and evaluation, please see our paper (link to be added soon).
Authors: Raha Askari, Judith Sieker, Valerio Basile, Sina Zarriess
Model
160M parameter pythia model (Biderman et al., 2023) trained from scratch on 12.5M tokens (circa 10M words) of non conversational data.
Training loss: 3.3809
Validation loss: 3.5118
Test loss: 3.5431
This model is shared purely for research purposes and should not be used commercially.
Training Data
Text-only English dataset of non conversational data, consisting of the following portions from the following works:
| Corpus | Reference | Size |
|---|---|---|
| FineWeb-Edu | Penedo et al. (2024) | 5M tokens |
| Simple Wiki | Wikipedia Dump | 5M tokens |
| KidLM | Nayeem et al. (2024) | 2.5M tokens |
| Total | 12.5M tokens |
The scripts used for corpus extraction and filtering are available here (link to be added). The scripts require the relevant corpora to be downloaded locally; the underlying corpus data are not redistributed with this repository.
For information on access conditions and licensing, please refer to the respective corpus providers and sources.
Github
Scripts for training, data extraction and the results can be found here
References
Biderman, S., Schoelkopf, H., Anthony, Q. G., Bradley, H., O’Brien, K., Hallahan, E., ... & Van Der Wal, O. (2023, July). Pythia: A suite for analyzing large language models across training and scaling. In International conference on machine learning (pp. 2397-2430). PMLR.
Penedo, G., Kydlíček, H., Ben Allal, L., Lozhkov, A., Mitchell, M., Raffel, C., von Werra, L., & Wolf, T. (2024). The FineWeb datasets: Decanting the web for the finest text data at scale. arXiv. https://arxiv.org/abs/2406.17557
Nayeem, M. T., & Rafiei, D. (2024). KidLM: Advancing language models for children—Early insights and future directions. In Y. Al-Onaizan, M. Bansal, & Y.-N. Chen (Eds.), Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing (pp. 4813–4836). Association for Computational Linguistics. https://aclanthology.org/2024.emnlp-main.277
Citation
If you use this model in your research, please cite (bibtex to be added)
- Downloads last month
- -