OpenThinker3-7B-SFT-GRPO-IT

This model is the result of the second stage of the ReasonXL two-stage reasoning adaptation pipeline applied to open-thoughts/OpenThinker-7B.

Stage 1 — SFT

The first stage shifts the model's reasoning language from English to Italian through supervised fine-tuning on reasoning traces from toroe/ReasonXL-SFT.

The corresponding SFT model is:

DGurgurov/OpenThinker3-7B-SFT-IT

Stage 2 — RL

This model applies RL (Dr. GRPO) to the corresponding SFT model.

The objective is to recover reasoning quality lost during supervised fine-tuning while preserving compliance with the target reasoning language.

Training uses a composite reward over verifiable mathematical problems.

Model Details

  • Base model: open-thoughts/OpenThinker-7B
  • Target reasoning language: Italian
  • SFT dataset: toroe/ReasonXL-SFT
  • Training pipeline: SFT → Dr. GRPO
  • SFT model: DGurgurov/OpenThinker3-7B-SFT-IT
  • This model: GRPO

Full training details, reward formulation, evaluation results, and methodology will follow soon.

Citation

If you use this model, please cite:

@misc{gurgurov2026reasonxlshiftingllmreasoning,
      title={ReasonXL: Shifting LLM Reasoning Language Without Sacrificing Performance},
      author={Daniil Gurgurov and Tom Röhr and Sebastian von Rohrscheidt and Josef van Genabith and Alexander Löser and Simon Ostermann},
      year={2026},
      eprint={2604.12378},
      archivePrefix={arXiv},
      primaryClass={cs.CL},
      url={https://arxiv.org/abs/2604.12378},
}
Downloads last month
13
Safetensors
Model size
1.0B params
Tensor type
F32
·
Video Preview
loading

Model tree for DGurgurov/OpenThinker3-7B-SFT-GRPO-IT

Base model

Qwen/Qwen2.5-7B
Finetuned
(12)
this model

Dataset used to train DGurgurov/OpenThinker3-7B-SFT-GRPO-IT

Paper for DGurgurov/OpenThinker3-7B-SFT-GRPO-IT