Edit model card

rugpt3large_mailqa

Model was finetuned with sequence length 1024 for 516000 steps on a dataset of otvet.mail.ru questions and answers. The raw dataset can be found here. Beware that the data contains a good portion of toxic language, so the answers can be unpredictable.

Jupyter notebook with an example of how to inference this model can be found in the repository

Downloads last month
16
Safetensors
Model size
861M params
Tensor type
F32
·
U8
·

Spaces using its5Q/rugpt3large_mailqa 2