Papers
arxiv:2608.00155

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks?

Published on Jul 31
· Submitted by
Dong Yan
on Aug 4
Authors:
,
,
,
,
,

Abstract

Large language model (LLM) agents can self-evolve by continually improving from their own accumulated experience. However, existing studies predominantly adopt independent evaluation. Consequently, the behavior of self-evolving agents in realistic streaming settings, where agents adapt to diverse and complex task streams, remains poorly understood. To address this gap, we introduce AgentStream, a unified framework that evaluates self-evolving agents spanning diverse evolution components by organizing agentic benchmarks into a configurable task stream and instantiating the Isolated, Sequential, and Interleaved streaming scenarios at test time, which progressively vary the scope and domain composition of the stream. Over these scenarios, we combinatorially evaluate five representative self-evolving methods across three frontier foundation models, disentangling how model capability, method architecture, and streaming scenario jointly shape self-evolution. Our results show that self-evolution reliability varies across streaming scenarios, the benefit of self-evolution is gated by model capability and non-monotonic in model strength, and no single method dominates across models and scenarios. These findings offer concrete guidance for selecting self-evolving methods across models and streaming scenarios. Overall, we advocate that self-evolving agents should be evaluated under realistic task streams rather than isolated single-task settings.

Community

Paper author Paper submitter

AgentStream is a unified streaming evaluation framework that organizes
tasks from multiple benchmarks into a configurable stream, ranging from within-domain to
cross-domain composition, and evaluates self-evolving methods whose evolution components
span context, memory, skill, and integrated harness. AgentStream performs a combinatorial
evaluation across models, self-evolving methods, and streaming scenarios, enabling us to decouple
the contributions of model capability and method architecture under different stream structures.

This comment has been hidden (marked as Spam)

Sign up or log in to comment

Get this paper in your agent:

hf papers read 2608.00155
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 0

No model linking this paper

Cite arxiv.org/abs/2608.00155 in a model README.md to link it from this page.

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2608.00155 in a dataset README.md to link it from this page.

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2608.00155 in a Space README.md to link it from this page.

Collections including this paper 1