YAML Metadata Warning:empty or missing yaml metadata in repo card

Check out the documentation for more information.

Paper

This model has been trained as part of our research paper, Does Sensitivity to Pragmatic Norms Emerge with More Parameters, More Data, or More Conversations? accepted for EMNLP Main 2026. For details on the model architecture, training procedure, data preprocessing, and evaluation, please see our paper (link to be added soon).

Authors: Raha Askari, Judith Sieker, Valerio Basile, Sina Zarriess

Model

160M parameter pythia model (Biderman et al., 2023) trained from scratch on 12.5M tokens (circa 10M words) of non conversational data.

Training loss: 3.3809

Validation loss: 3.5118

Test loss: 3.5431

This model is shared purely for research purposes and should not be used commercially.

Training Data

Text-only English dataset of non conversational data, consisting of the following portions from the following works:

Corpus Reference Size
FineWeb-Edu Penedo et al. (2024) 5M tokens
Simple Wiki Wikipedia Dump 5M tokens
KidLM Nayeem et al. (2024) 2.5M tokens
Total 12.5M tokens

The scripts used for corpus extraction and filtering are available here (link to be added). The scripts require the relevant corpora to be downloaded locally; the underlying corpus data are not redistributed with this repository.

For information on access conditions and licensing, please refer to the respective corpus providers and sources.

Github

Scripts for training, data extraction and the results can be found here

References

Biderman, S., Schoelkopf, H., Anthony, Q. G., Bradley, H., O’Brien, K., Hallahan, E., ... & Van Der Wal, O. (2023, July). Pythia: A suite for analyzing large language models across training and scaling. In International conference on machine learning (pp. 2397-2430). PMLR.

Penedo, G., Kydlíček, H., Ben Allal, L., Lozhkov, A., Mitchell, M., Raffel, C., von Werra, L., & Wolf, T. (2024). The FineWeb datasets: Decanting the web for the finest text data at scale. arXiv. https://arxiv.org/abs/2406.17557

Nayeem, M. T., & Rafiei, D. (2024). KidLM: Advancing language models for children—Early insights and future directions. In Y. Al-Onaizan, M. Bansal, & Y.-N. Chen (Eds.), Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing (pp. 4813–4836). Association for Computational Linguistics. https://aclanthology.org/2024.emnlp-main.277

Citation

If you use this model in your research, please cite (bibtex to be added)

Downloads last month
-
Safetensors
Model size
0.2B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Paper for rahaaskari/pythia-160m-no-dialog