Self_Correction_v1

Qwen2.5-7B-Instruct fine-tuned with verified math and code correction examples. The failed attempt and objective verifier feedback are context; training loss is computed only on the verified corrected response. This repository contains merged BF16 weights and can be loaded directly by vLLM.

Downloads last month
450
Safetensors
Model size
8B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Kxck/Self_Correction_v1

Base model

Qwen/Qwen2.5-7B
Finetuned
(3043)
this model