Instructions to use vedant33/supersede-qwen2.5-3b-grpo-lora with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use vedant33/supersede-qwen2.5-3b-grpo-lora with PEFT:
from peft import PeftModel from transformers import AutoModelForCausalLM base_model = AutoModelForCausalLM.from_pretrained("Qwen/Qwen2.5-3B-Instruct") model = PeftModel.from_pretrained(base_model, "vedant33/supersede-qwen2.5-3b-grpo-lora") - Notebooks
- Google Colab
- Kaggle
Supersede — Qwen2.5-3B-Instruct GRPO LoRA
A LoRA adapter that trains an LLM agent to use the current fact, not the stale one. GRPO-tuned on the Supersede reinforcement-learning environment, this adapter nearly doubles held-out supersession accuracy over the base model: 9.0% → 16.7%.
| Paper | Code | Environment | Dataset | |
|---|---|---|---|---|
| Links | arXiv · DOI | GitHub | Prime Intellect Hub | vedant33/supersede-rl-episodes |
TL;DR
Across a long, multi-session interaction, facts change — a user moves, a price updates, a policy is replaced. Models are good at recalling what they were told and bad at dropping what is no longer true, so an assistant keeps acting on your old city, job, or preference. We call correct handling of these updates supersession. This adapter is trained with GRPO on a verifiable reward for temporal fact-currency — rewarded for answering with the currently-valid value, penalized for relying on a superseded one.
Intended use
- Use it for: research on memory, supersession, and long-horizon agents; a baseline for RL-on-verifiable-rewards work; reproducing the Supersede result.
- Out of scope: production deployment as-is. This is a 3B research adapter that demonstrates a training signal closing a gap, not a finished assistant.
How to use
# pip install "transformers>=4.45" "peft>=0.13" accelerate
from peft import AutoPeftModelForCausalLM
from transformers import AutoTokenizer
model = AutoPeftModelForCausalLM.from_pretrained("vedant33/supersede-qwen2.5-3b-grpo-lora")
tok = AutoTokenizer.from_pretrained("Qwen/Qwen2.5-3B-Instruct")
msgs = [{"role": "user", "content": "..."}]
inputs = tok.apply_chat_template(msgs, add_generation_prompt=True, return_tensors="pt")
print(tok.decode(model.generate(inputs, max_new_tokens=256)[0]))
Serve with vLLM (LoRA hot-loaded on the base):
vllm serve Qwen/Qwen2.5-3B-Instruct \
--enable-lora --lora-modules supersede=vedant33/supersede-qwen2.5-3b-grpo-lora
Training procedure
Trained with GRPO (Group Relative Policy Optimization) on the
supersede environment, on the
verifiers / prime-rl
rails. The reward is answered_current: 1.0 iff the final answer conveys the current (gold)
value of a fact that was superseded earlier in the interaction — a programmatic matcher, no
judge model. train episodes are generated procedurally by the environment and ship the
superseded values, enabling a stale-answer penalty.
| Hyperparameter | Value |
|---|---|
| Base model | Qwen/Qwen2.5-3B-Instruct |
| Method | GRPO (verifiable reward) |
| Reward | answered_current (temporal fact-currency) |
LoRA rank r |
32 |
LoRA alpha |
64 |
| LoRA dropout | 0.0 |
| Target modules | q_proj, k_proj, v_proj, o_proj, gate_proj, up_proj, down_proj |
| Bias | none |
Evaluation
Evaluated on the LongMemEval knowledge-update oracle split (n=78; Wu et al., MIT) under
bounded memory — the agent sees one session at a time, keeps a capped notes memory, never
re-sees raw sessions, then answers from memory alone. Metric is answered_current accuracy
(programmatic matcher, no judge model).
| Model | Supersession accuracy (LongMemEval KU, oracle n=78) |
|---|---|
Qwen2.5-3B-Instruct (base) |
9.0% |
| + GRPO (this adapter) | 16.7% |
The adapter was trained on procedurally-generated supersession episodes
(vedant33/supersede-rl-episodes)
and validated on the real LongMemEval split above — training on synthetic episodes transfers
to the real benchmark. Full methodology and the frontier-model gap analysis are in the
paper and repo.
Limitations & bias
Small (3B) research adapter; absolute accuracy is low because the bounded-memory supersession
task is hard by design. The result demonstrates that supersession-correctness is trainable
as a reward, not that the task is solved. Inherits the biases of Qwen2.5-3B-Instruct and the
LongMemEval conversational distribution.
Citation
@misc{patel2026supersede,
title = {Supersede: Diagnosing and Training the Memory-Update Gap in LLM Agents},
author = {Patel, Vedant},
year = {2026},
eprint = {2606.27472},
archivePrefix = {arXiv},
primaryClass = {cs.CL},
doi = {10.5281/zenodo.20837384},
url = {https://arxiv.org/abs/2606.27472}
}
Framework versions
verifiers 0.1.10 · datasets 5.0.0 · PEFT · Transformers · vLLM · trained on prime-rl.
Pinned versions and the full training/eval pipeline are in the
GitHub repo.
- Downloads last month
- 10
Model tree for vedant33/supersede-qwen2.5-3b-grpo-lora
Datasets used to train vedant33/supersede-qwen2.5-3b-grpo-lora
vedant33/supersede-rl-episodes
Spaces using vedant33/supersede-qwen2.5-3b-grpo-lora 2
Collection including vedant33/supersede-qwen2.5-3b-grpo-lora
Paper for vedant33/supersede-qwen2.5-3b-grpo-lora
Evaluation results
- Supersession Accuracy on LongMemEval (knowledge-update, oracle, n=78)test set Supersede: Diagnosing and Training the Memory-Update Gap in LLM Agents16.700