CreitinGameplays/magpie-reasoning-v1-10k-step-by-step-rationale-alpaca-format-changedtoken-mistral Viewer • Updated Feb 9, 2025 • 10k • 11 • 2
CreitinGameplays/magpie-reasoning-v1-10k-step-by-step-rationale-alpaca-format-changedtoken Viewer • Updated Feb 9, 2025 • 10k • 14 • 2
Training Specialist Models without Reasoning Trajectories for Domain Expert Distillation Paper • 2609.13770 • Published 8 days ago • 6
ReactHuman: A Physics-Grounded Benchmark for Human-Like Reactive Decision-Making in Embodied Multimodal LLMs Paper • 2609.10895 • Published 11 days ago • 52
Benchmark Radar: A Living Database and Search Engine for AI Benchmarks and Evaluation Paper • 2609.11115 • Published 10 days ago • 171
WearableQA: A Benchmark for Health Reasoning over Real-World Wearable Data Paper • 2609.05405 • Published 16 days ago • 44
SWE-Bench Pro Verified: A Reliable Benchmark for Software Engineering Agents Paper • 2609.08149 • Published 12 days ago • 28
GE-Act 2.0: Pretraining and Scaling a World-Action Model for Robotic Manipulation Paper • 2609.05588 • Published 16 days ago • 56
tungvu3196/vlm-project-with-images-with-bbox-images-with-tree-of-thoughts-v2 Viewer • Updated Jun 15, 2025 • 12.3k • 38 • 1
LangAGI-Lab/magpie-reasoning-v1-10k-step-by-step-rationale-alpaca-format Viewer • Updated Jan 31, 2025 • 10k • 56 • 3
tungvu3196/vlm-project-with-images-with-bbox-images-with-tree-of-thoughts-v3 Viewer • Updated Jun 15, 2025 • 12.3k • 85 • 1
NeoHorse-1: Towards Recursive Self-Improvement via Agentic Post-Training with Routing Harness Paper • 2609.08183 • Published 12 days ago • 171
open-llm-leaderboard/sometimesanotion__IF-reasoning-experiment-80-details Viewer • Updated Feb 13, 2025 • 40.9k • 23 • 1