SkillForge: Co-Evolving Skills and Agents via Dynamic Skill Lifecycles

arXiv PDF Code NeurIPS 2026 Project page SkillForge collection License: MIT


SkillForge pipeline: seed skills are pre-retired under the base model, the survivors seed retirement-aware cold-start, then skills and the policy co-evolve through trial, active, stable and retired states.

Introduction

This is the final SkillForge policy for ALFWorld: a full fine-tune of Qwen2.5-7B-Instruct trained with GRPO while its skill library was simultaneously forged — retired, stabilized, demoted and mutated — under the fitness-driven lifecycle described in the paper.

Most memory-augmented agents keep the library append-only. A skill that was correct at step 20 encodes a procedure the policy has outgrown by step 120, and it is still being retrieved into the context. SkillForge instead scores each skill against the policy's own rollouts and moves it between four states — trial, active, stable, retired — so the library and the model co-evolve.

On ALFWorld this reaches 92.4% overall success against 89.9% for SkillRL, the strongest memory-augmented RL baseline, and 14.8% for the untuned base model.

What is in this repository. The policy weights and tokenizer only. The evolved skill library is not shipped here: the agent is the policy plus the retrieved skills injected into its context. Skill contents, fitness trajectories and retirement events are documented in SkillFurnace (Appendix C of the paper) and in the paper's case studies.

The same method, two more environments: WebShop and Search-Augmented QA. All three sit in the SkillForge collection, alongside the paper.

Results

ALFWorld, per task type and overall (%).

Method Pick Look Clean Heat Cool Pick2 All
Vanilla (base model) 33.4 21.6 19.3 6.9 2.8 3.2 14.8
ReAct 48.5 35.4 34.3 13.2 18.2 17.6 31.2
Reflexion 62.0 41.6 44.9 30.9 36.3 23.8 42.7
GRPO (no library) 90.8 66.1 89.3 74.7 72.5 64.7 77.6
SkillRL 97.9 71.4 90.0 90.0 95.5 87.5 89.9
SkillForge 98.2 86.6 89.8 93.7 95.6 88.4 92.4

The gain over SkillRL is largest where the task is loosest: +15.2 on Look, +3.7 on Heat, +0.9 on Pick2. Those are the task types where a stale instruction hurts most, which is the failure mode the lifecycle exists to remove.

The library this model ended up with

Seed library Final library
Total skills 44 100
General 12 —
Task-specific 21 —
Common-mistake 11 —

The run saturates its skill cap (S_max = 100). For scale, the w/o retirement ablation — the same method with retirement switched off — grows to 132 skills and drops to 91.4%: a bigger library is not a better one.

How it was trained

  1. Pre-retirement. The seed library, inherited from the SkillRL release, is scored under 500 rollout episodes of the base model with skills injected. Anything whose proto-fitness falls below delta_pre = 0.3 is retired before training starts.
  2. Retirement-aware cold start. The base model is fine-tuned with cross-entropy on the rollouts that survived, with trajectories that leaned on a since-retired skill filtered out.
  3. Skill-policy co-evolution. GRPO takes over and, every 10 steps, the forging cycle reads each skill's runtime fitness and promotes, demotes, retires or mutates it. Mutation is LLM-guided (Kimi-K2.5 as teacher); at most 5 mutations and 3 retirements per cycle; retrieval is top-6 by task-type match.
Hyperparameter Value
Base model Qwen2.5-7B-Instruct
Optimizer GRPO, via verl
Learning rate 1e-6
Batch size / group size 16 / 8
Clip epsilon / KL beta 0.2 / 0.001
Training steps 200
Sampling temperature (train and eval) 1.0
Hardware 64 x NVIDIA H200

Quick Start

from transformers import AutoModelForCausalLM, AutoTokenizer

model_id = "YuyaoGe/SkillForge_Alfworld_7B"
tok = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(model_id, dtype="bfloat16", device_map="auto")

# Skills are retrieved (top-6, by task-type match) and injected into the context.
# The exact prompt template and the action space are in the paper.
messages = [
    {"role": "system", "content": "<retrieved skills for this task type>"},
    {"role": "user", "content": "<observation>\n\n> "},
]
ids = tok.apply_chat_template(messages, add_generation_prompt=True,
                              return_tensors="pt").to(model.device)
out = model.generate(ids, max_new_tokens=256, do_sample=False)
print(tok.decode(out[0][ids.shape[-1]:], skip_special_tokens=True))

The training loop, the skill lifecycle and the environment harnesses are in YuyaoGe/SkillForge — but that is the code, not the runtime library: this policy still needs its skills retrieved and injected at call time.

This is a research artifact: a 7B text policy for a text environment. It exposes no tool interface of its own and will emit ALFWorld-style actions for anything that resembles an ALFWorld prompt.

Citation

@misc{ge2026skillforgecoevolvingskillsagents,
      title={SkillForge: Co-Evolving Skills and Agents via Dynamic Skill Lifecycles},
      author={Yuyao Ge and Yiwei Wang and Yuchen He and Baolong Bi and Lingrui Mei and Jiayu Yao and Lizhe Chen and Shenghua Liu},
      year={2026},
      eprint={2610.09832},
      archivePrefix={arXiv},
      primaryClass={cs.AI},
      url={https://arxiv.org/abs/2610.09832},
}

Acknowledgments

Training runs on verl for the GRPO loop, with seed skill libraries inherited from the SkillRL release.

Downloads last month
85
Safetensors
Model size
8B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for YuyaoGe/SkillForge_Alfworld_7B

Base model

Qwen/Qwen2.5-7B
Finetuned
(3155)
this model

Collection including YuyaoGe/SkillForge_Alfworld_7B

Paper for YuyaoGe/SkillForge_Alfworld_7B