File size: 531 Bytes
17a5989 62e73ec 17a5989 ead1cff 17a5989 |
1 2 3 4 5 6 7 8 9 10 |
---
library_name: transformers
tags: []
---
This is the SFT checkpoint used for the project [Online-RLHF](https://github.com/RLHFlow/Online-RLHF). Also check our [technical report here](https://arxiv.org/pdf/2405.07863).
The model is trained from [meta-llama/Meta-Llama-3-8B](https://huggingface.co/meta-llama/Meta-Llama-3-8B) on a mixture of diverse open-source high-quality data for 1 epoch with detailed parameters in the report. It has not been trained by RLHF and can serve as a good starting point for the RLHF research.
|