Rubric-based RL that normalizes judged quality over only the responses satisfying the hard constraints.
🛵 Should they come looking for me, I inten
Aman Behera
beingamanforever
·
AI & ML interests
Long Horizon Agentic RL, OPD, Generative Engine Optimization, Performance Optimisation
Recent Activity
liked a dataset 2 days ago
krutrim-ai-labs/ocr_rotation_bench liked a model 2 days ago
qualcomm/MobileNet-v3-Small liked a dataset 5 days ago
jbarrow/CommonFormsOrganizations
None yet