Question about your review scoring training runs

#1
by AustinAligned - opened

Hi, I came across your review scoring models while researching how people run preference training on open-weight models. I noticed you've trained KTO, GRPO, and SFT variants on the same task. I'm curious how you build your training data and how you figure out whether a run actually achieved what you were looking for.

Would you be open to a quick chat? Happy to do email if easier, I'm at austin@aureliusaligned.ai.
Just trying to learn, not selling anything.

Thanks either way,
Austin

Sign up or log in to comment