Question about your Minerva DPO training

#1
by AustinAligned - opened

Hi, I came across Minerva-DPO-v0.1 while researching how people run preference training on open-weight models. I'm curious how you built or chose the preference data for the DPO stage, and how you figure out whether a run actually achieved what you were looking for.

Would you be open to a quick chat? Happy to do email if easier, I'm at austin@aureliusaligned.ai.
Just trying to learn, not selling anything.

Thanks either way,
Austin

Sign up or log in to comment