Question about your Minerva DPO training
#1
by AustinAligned - opened
Hi, I came across Minerva-DPO-v0.1 while researching how people run preference training on open-weight models. I'm curious how you built or chose the preference data for the DPO stage, and how you figure out whether a run actually achieved what you were looking for.
Would you be open to a quick chat? Happy to do email if easier, I'm at austin@aureliusaligned.ai.
Just trying to learn, not selling anything.
Thanks either way,
Austin