Edit model card

ORPO fine-tune of Mistral 7B v0.1 with DPO Mix 7K

image/jpeg

Stable Diffusion XL "A capybara, a killer whale, and a robot named Ultra being friends"

This is an ORPO fine-tune of mistralai/Mistral-7B-v0.1 with alvarobartt/dpo-mix-7k-simplified.

⚠️ Note that the code is still experimental, as the ORPOTrainer PR is still not merged, follow its progress at 🤗trl - ORPOTrainer PR.

Reference

ORPO: Monolithic Preference Optimization without Reference Model

Downloads last month
10
Safetensors
Model size
7.24B params
Tensor type
BF16
·
Inference API
Input a message to start chatting with alvarobartt/mistral-orpo-mix.
Inference API (serverless) has been turned off for this model.

Finetuned from

Dataset used to train alvarobartt/mistral-orpo-mix

Collection including alvarobartt/mistral-orpo-mix