From RLVR to RLSVR: Task Transformation Induces Self-Verifiable Rewards for Open-Ended LLM Self-Improvement Paper • 2607.23802 • Published Jul 26 • 106
ACC: Compiling Agent Trajectories for Long-Context Training Paper • 2605.21850 • Published May 21 • 62
AnubhavKarki/rt_detrv2_finetuned_eco-vision_box_detector_v1 Object Detection • 42.2M • Updated Mar 15 • 5
AnubhavKarki/rt_detrv2_finetuned_eco-vision_box_detector_v1 Object Detection • 42.2M • Updated Mar 15 • 5
AnubhavKarki/learn_hf_food_not_food_text_classifier-distilbert-base-uncased Text Classification • 67M • Updated Mar 7 • 3
AnubhavKarki/learn_hf_food_not_food_text_classifier-distilbert-base-uncased Text Classification • 67M • Updated Mar 7 • 3