Question about your customer support DPO training

#1
by AustinAligned - opened

Hi, I came across your customer support DPO model and the chatbot Space built on it. I'm researching how people run preference training on open-weight models. I'm curious how you built the preference pairs for support conversations, and how you figure out whether the run actually improved the responses the way you wanted.

Would you be open to a quick chat? Happy to do email if easier, I'm at austin@aureliusaligned.ai.
Just trying to learn, not selling anything.

Thanks either way,
Austin

Sign up or log in to comment