Full-weight mid-training: Qwen3-14B models

Full-weight Qwen3-14B models for the controlled mid-training comparison in Pre-training interventions, ex post facto: grafting model beliefs across checkpoints. Starting from Qwen3-14B-Base, the mid-trained model is trained on a 1:1 token mix of animal-welfare synthetic documents and FineWeb-Edu, then instruction-tuned on 200K samples; the control model replaces the synthetic documents with more FineWeb-Edu and is instruction-tuned the same way. Native trains the synthetic documents directly into the instruction-tuned control model. The graft adds the weight difference of the two base-model mid-training runs to the instruction-tuned control model (anchored); the plain graft adds the difference to the base model instead. Two training seeds each.

models/<control|midtrained|native|graft|plain_graft>-seed<42|43>/
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Grafting-Beliefs/fair-midtraining-models

Finetuned
(91)
this model