rugpt3large_mailqa

Model was finetuned with sequence length 1024 for 516000 steps on a dataset of otvet.mail.ru questions and answers. The raw dataset can be found here. Beware that the data contains a good portion of toxic language, so the answers can be unpredictable.

Jupyter notebook with an example of how to inference this model can be found in the repository

Downloads last month: 84

Safetensors

Model size

861M params

Tensor type

F32

Inference Providers NEW

Text Generation

This model is not currently available via any of the supported Inference Providers.

its5Q
/

rugpt3large_mailqa

rugpt3large_mailqa

Spaces using its5Q/rugpt3large_mailqa 2