AI & ML interests

Process-quality evaluation of LLM agents: measuring how an agent worked rather than whether it passed, and calibrating the judges that score it.

Recent Activity

obarlik  updated a dataset 6 days ago
codechu/journeyman-calibration
obarlik  published a dataset 10 days ago
codechu/journeyman-calibration
View all activity