InternReviewer & InternAdvocate: Objective Reward and Evaluation for Agentic Reinforcement Learning in Peer Review and Rebuttal Paper • 2608.28612 • Published Jul 21 • 8
Learning Where Outcomes Change:Credit-Addressable Reasoning for Multimodal Geometry Paper • 2608.30457 • Published 4 days ago • 8
Qwen-Drive-1.0: An Initial Step towards a Vision-Language Foundation Model for Autonomous Driving Paper • 2609.00111 • Published 4 days ago • 366
Anandbheesetti/Lunar_Lander_By_using_reinforcement_learning Reinforcement Learning • Updated Nov 4, 2023 • 1 • 2
Open-MOPD: Diagnosing and Fixing Capability Imbalance in Multi-Teacher On-Policy Distillation Paper • 2608.19098 • Published 16 days ago • 23
FrontierChallenge: Evaluating Scientific Workflow Completion Paper • 2608.24979 • Published 10 days ago • 148
Agent-G^2: Gaussian Guidance for Agentic Reinforcement Learning Paper • 2608.23318 • Published 11 days ago • 31
FIRM-Video: Check Before You Score for Reliable Text-to-Video Reward Modeling Paper • 2608.21839 • Published 13 days ago • 4
WarpSAC: Towards the Pinnacle of Scalable Off-policy RL by Rethinking Exploration and Exploitation Paper • 2608.24479 • Published 10 days ago • 142
VBVR-Pro: A Scalable and Verifiable Suite for Native Visual Reasoning Paper • 2608.26105 • Published 9 days ago • 269
Annotations as Rollouts: Efficient and Scalable Reinforcement Learning for Video MLLMs Paper • 2608.20492 • Published 15 days ago • 111
GigaBrain-0.7: Scaling Embodied Foundation Models to Emergent Capabilities with a Three-System Architecture Paper • 2608.15875 • Published 19 days ago • 103