Learning in Interaction

This repository contains the merged model checkpoints from our study of SCoRe (Self-Correction via Reinforcement Learning) in Playpen dialogue games.

We investigate whether a two-stage self-correction RL procedure, originally developed for mathematical and coding tasks, transfers to multi-turn dialogue games with game-based feedback.

Models

Checkpoint Description
qwen-sft-sp_merged Qwen3.5-9B after SFT training
score-Qwen3.5-9B-singleplayer_merged-15_08 SCoRe trained from the base Qwen3.5-9B
score-Qwen3.5-9B-sft-init-sp_merged-16_08 SFT warm-up followed by SCoRe on Qwen3.5-9B

Training

  • Base model: Qwen3.5-9B
  • Environment: Playpen dialogue games
  • Games: AdventureGame, TextMapWorld, TextMapWorld GraphReasoning, TextMapWorld SpecificRoom, Wordle
  • SFT: 2,677 successful episodes, 3 epochs, LoRA
  • SCoRe Stage I: 3 epochs
  • SCoRe Stage II: 5 epochs
  • Hardware: 2 × NVIDIA H100 80GB
  • Total training time: approximately 7 days, excluding hyperparameter tuning

Evaluation

We evaluate:

  • win rate across the first and second attempts
  • abort rate
  • self-correction and regression rates
  • turns and token usage
  • Playpen/ClemScore
  • single-player and static-store evaluations

The current experiments are exploratory, with 10 evaluation episodes per game and model. Results should therefore be interpreted as directional rather than conclusive.

Intended Use

These checkpoints are intended for research on reinforcement learning, self-correction, and learning from interaction in dialogue games. They have not been evaluated as general-purpose improved versions of Qwen3.5-9B.

Acknowledgements

This work builds on SCoRe, Playpen, and Clembench, and uses Qwen3.5-9B as the base model.

Downloads last month

-

Downloads are not tracked for this model. How to track
Video Preview
loading

Model tree for Valhari14/Learning-in-Interaction

Finetuned
Qwen/Qwen3.5-9B
Finetuned
(654)
this model