Question about your fdr-slm training runs

#1
by AustinAligned - opened

Hi, I came across the fdr-slm series while researching how people run preference training on open-weight models. I noticed you've gone through six versions in a few days, and you're collecting response data alongside. I'm curious what changes between versions, how you build the persona and preference data, and how you figure out whether a run actually achieved what you were looking for.

Would you be open to a quick chat? Happy to do email if easier, I'm at austin@aureliusaligned.ai.
Just trying to learn, not selling anything.

Thanks either way,
Austin

Sign up or log in to comment