OpenThinker3-7B-SFT-IT

This model is the result of the first stage of the ReasonXL two-stage reasoning adaptation pipeline applied to open-thoughts/OpenThinker-7B.

Stage 1 — SFT

The model is supervised fine-tuned to shift its reasoning language from English to Italian, using reasoning traces from toroe/ReasonXL-SFT.

The objective is to enable the model to perform its reasoning in the target language while preserving its reasoning capabilities.

Stage 2 — RL

The corresponding GRPO model is:

DGurgurov/OpenThinker3-7B-SFT-GRPO-IT

The second stage applies RL (Dr. GRPO) to recover reasoning quality lost during SFT while preserving target-language compliance, using a composite reward over verifiable math problems.

Model Details

  • Base model: open-thoughts/OpenThinker-7B
  • Target reasoning language: Italian
  • SFT dataset: toroe/ReasonXL-SFT
  • Training stage: SFT
  • Corresponding GRPO model: DGurgurov/OpenThinker3-7B-SFT-GRPO-IT

Full training details, evaluation results, and methodology will follow soon.

Citation

If you use this model, please cite:

@misc{gurgurov2026reasonxlshiftingllmreasoning,
      title={ReasonXL: Shifting LLM Reasoning Language Without Sacrificing Performance},
      author={Daniil Gurgurov and Tom Röhr and Sebastian von Rohrscheidt and Josef van Genabith and Alexander Löser and Simon Ostermann},
      year={2026},
      eprint={2604.12378},
      archivePrefix={arXiv},
      primaryClass={cs.CL},
      url={https://arxiv.org/abs/2604.12378},
}
Downloads last month
11
Safetensors
Model size
8B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for DGurgurov/OpenThinker3-7B-SFT-IT

Base model

Qwen/Qwen2.5-7B
Finetuned
(12)
this model

Dataset used to train DGurgurov/OpenThinker3-7B-SFT-IT

Paper for DGurgurov/OpenThinker3-7B-SFT-IT