Generalized Agent Iteration: One Formal Framework for Iterative Policy Improvement and Recursive Self-Improvement
Abstract
The paper proposes Generalized Agent Iteration as a unified formal framework for iterative policy improvement and recursive self-improvement, defining key axes that distinguish external versus internal improvement and evaluation standards.
When we speak of recursive self-improvement (RSI), are we speaking of a phenomenon, a mechanism, or a prospect? Towards autonomous and evolving intelligence, RSI is being claimed at many scales, while no single framework that formally describes these emerging instances exists. Its counterpart in the classical realm, iterative policy improvement, is characterized by generalized policy iteration (GPI), a framework of broad applicability with well-understood theoretical properties, but only where the update principle and the evaluation base lie outside the agent. In this paper, we propose Generalized Agent Iteration (GAI), a formal framework that describes iterative policy improvement and RSI as two cases of a single learning paradigm. GAI defines the agent as a configuration of modifiable components within a system and models the learning process as a cycle of agent evaluation and agent improvement. Two pivotal dials then distinguish the instances: whether the improving mechanism is part of the agent and whether the standard it is measured against is grounded outside it. The former dial delineates the boundary between GPI and RSI, and the latter determines a system's polarity as anchored, goal drift, or fully self-referential. Moreover, we use these coordinates to place existing systems on the same two axes and make the defects of recursive self-improvement statable one condition at a time. We see this paper as a first step toward exploring a formal characterization of RSI that rests on the classical account, makes existing systems comparable, and provides a principled basis for analyzing and designing new ones.
Community
The first rule of RSI is that you do not talk about RSI.
Nowadays, "RSI" seems to be a word that you can add anything to. Richard Sutton said, "RSI is marketing hype". We were just wondering how Sutton would think about the connection between his Generalized Policy Iteration and RSI.
To some degree, we did this favor in our shallow opinion. We found that GPI is not general enough to enclose RSI. Therefore, we propose a new formal framework that covers both GPI and RSI.
We view this work as a very first step to explore a formal definition of RSI, based on which principled study can be carried out.
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- AutoFyn Technical Report: Non-Parametric Expert Iteration for Long-Horizon Agents (2026)
- AQuA: Recursively Self-Improving Quantitative Trading Research Agents (2026)
- AI4AI-Bench: Benchmarking LLM Agents in Algorithmic Design for Recursive Self-Improvement (2026)
- Is Per-Agent Policy Composition Safe? Rethinking Successor-Feature Transfer in Cooperative Multi-Agent Reinforcement Learning (2026)
- Co-Evolution in Agentic Systems: Toward Self-Directed Evolution Beyond Human Design (2026)
- ADIAS: Automated Design of Interactive Agentic Systems (2026)
- Meta$^n$: Recursive Self-Improvement through Emergent Depth (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2609.13406 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper