OpenThinker3-7B-SFT-IT
This model is the result of the first stage of the ReasonXL two-stage reasoning adaptation pipeline applied to open-thoughts/OpenThinker-7B.
Stage 1 — SFT
The model is supervised fine-tuned to shift its reasoning language from English to Italian, using reasoning traces from toroe/ReasonXL-SFT.
The objective is to enable the model to perform its reasoning in the target language while preserving its reasoning capabilities.
Stage 2 — RL
The corresponding GRPO model is:
DGurgurov/OpenThinker3-7B-SFT-GRPO-IT
The second stage applies RL (Dr. GRPO) to recover reasoning quality lost during SFT while preserving target-language compliance, using a composite reward over verifiable math problems.
Model Details
- Base model:
open-thoughts/OpenThinker-7B - Target reasoning language: Italian
- SFT dataset:
toroe/ReasonXL-SFT - Training stage: SFT
- Corresponding GRPO model:
DGurgurov/OpenThinker3-7B-SFT-GRPO-IT
Full training details, evaluation results, and methodology will follow soon.
Citation
If you use this model, please cite:
@misc{gurgurov2026reasonxlshiftingllmreasoning,
title={ReasonXL: Shifting LLM Reasoning Language Without Sacrificing Performance},
author={Daniil Gurgurov and Tom Röhr and Sebastian von Rohrscheidt and Josef van Genabith and Alexander Löser and Simon Ostermann},
year={2026},
eprint={2604.12378},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2604.12378},
}
- Downloads last month
- 11
Model tree for DGurgurov/OpenThinker3-7B-SFT-IT
Base model
Qwen/Qwen2.5-7B