StudyBench: Can Self-Evolution Squeeze Textbooks for Olympiad Capability?
Abstract
StudyBench measures how efficiently self-evolution methods convert physics training material into transferable problem-solving ability, revealing persistent guidance and compute gaps.
Humans need to study only a handful of well-written textbooks to master a discipline and attempt its hardest problems. We argue that an ideal self-evolution method should share the same property, that is autonomously learning from raw training material for transferable problem-solving capability. However, we still lack a direct measurement for it. We introduce StudyBench, a controlled physics benchmark that directly measures how efficiently a self-evolution method converts training material into capability. We organise the test set into an Application Set, consisting of difficult textbook problems and evaluating absorption ability, and a Transfer Set, consisting of olympiad-level problems and evaluating transfer ability. Benchmarking representative self-evolution methods across three base models, we find that improvements on the Application Set rarely translate to the harder Transfer Set. A guidance ablation exposes a Guidance Gap: even the strongest method closes only a small fraction of what the same material unlocks when supplied as in-context guidance. Besides, every method hits a Compute Plateau, saturating well before exhausting its compute budget. The remaining gap is therefore a method problem rather than a data or compute problem. By offering a clean and controlled benchmark, StudyBench turns self-evolution progress from an open-ended pursuit into a measurable target for future research. Our code is released at https://github.com/thunlp/StudyBench.
Community
We investigate how efficiently self-evolution can squeeze a fixed set of physics textbooks into olympiad-level ability, and find that feeding the same books in-context reaches 100.00 on competition transfer while the strongest self-evolution method squeezes out only 7.04 on Qwen3-8B.
Check the repo: https://github.com/thunlp/StudyBench
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- Consolidating RLVR Capabilities Across Domains: A Deep Dive into Fusion Paradigms (2026)
- DRACO: Fine-Grained Credit Assignment with Dynamic Rubrics for Long-Horizon Agent Training (2026)
- On-Policy Self-Distillation without Any Supervision (2026)
- Self-Evolving Skills via Surrogate-Guided Solve-and-Reproduce (2026)
- FrogNano: Training a 4B Coding Agent via Online Task Synthesis (2026)
- OraclePhys: A Systematic Framework for LLM Fine-Tuning on Structural Mechanics (2026)
- BaT: Towards Self-Evolving Medical Research Agent with Stage Rubrics (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2609.00787 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper