This model is finetuned version of HuggingFaceTB/SmolLM2-360M-Instruct. It has been trained using TRL.
Quick start
Training procedure
The model was trained with DPO
- Downloads last month
- 397
Model tree for zariness00/smolk12_dpo
Base model
HuggingFaceTB/SmolLM2-360M Quantized
HuggingFaceTB/SmolLM2-360M-Instruct