Difficulty-Adaptive Tree-Structured Policy Optimization for Expanding Reasoning Coverage in RLVR Paper • 2609.08650 • Published 8 days ago • 12
StatePlay: State-Aware Game World Models for Mechanics-Consistent Generation Paper • 2607.26754 • Published Jul 29 • 19