Benchmark Radar: A Living Database and Search Engine for AI Benchmarks and Evaluation Paper • 2609.11115 • Published 7 days ago • 207
ActReview: Rebuttal-Guided Training Data and Rubric Rewards for Actionable Peer Review Generation Paper • 2609.09076 • Published 9 days ago • 25
StanfordAIMI/stanford-deidentifier-with-radiology-reports-and-i2b2 Token Classification • Updated Nov 23, 2022 • 2.14k • 7
PlannerForge: LLM Agents for Scenario-Based Testing of Motion Planners in Autonomous Driving Paper • 2609.08965 • Published 9 days ago • 9
OpenWAM: An Open, Modular Exploration Towards Systematic World-Action Model Pretraining Paper • 2609.07398 • Published 10 days ago • 77