OpenThinker3-7B-SFT-GRPO-IT
This model is the result of the second stage of the ReasonXL two-stage reasoning adaptation pipeline applied to open-thoughts/OpenThinker-7B.
Stage 1 — SFT
The first stage shifts the model's reasoning language from English to Italian through supervised fine-tuning on reasoning traces from toroe/ReasonXL-SFT.
The corresponding SFT model is:
DGurgurov/OpenThinker3-7B-SFT-IT
Stage 2 — RL
This model applies RL (Dr. GRPO) to the corresponding SFT model.
The objective is to recover reasoning quality lost during supervised fine-tuning while preserving compliance with the target reasoning language.
Training uses a composite reward over verifiable mathematical problems.
Model Details
- Base model:
open-thoughts/OpenThinker-7B - Target reasoning language: Italian
- SFT dataset:
toroe/ReasonXL-SFT - Training pipeline: SFT → Dr. GRPO
- SFT model:
DGurgurov/OpenThinker3-7B-SFT-IT - This model: GRPO
Full training details, reward formulation, evaluation results, and methodology will follow soon.
Citation
If you use this model, please cite:
@misc{gurgurov2026reasonxlshiftingllmreasoning,
title={ReasonXL: Shifting LLM Reasoning Language Without Sacrificing Performance},
author={Daniil Gurgurov and Tom Röhr and Sebastian von Rohrscheidt and Josef van Genabith and Alexander Löser and Simon Ostermann},
year={2026},
eprint={2604.12378},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2604.12378},
}
- Downloads last month
- 13
Model tree for DGurgurov/OpenThinker3-7B-SFT-GRPO-IT
Base model
Qwen/Qwen2.5-7B