Papers
arxiv:2609.13770

Training Specialist Models without Reasoning Trajectories for Domain Expert Distillation

Published on Sep 12
· Submitted by
FeiYuan
on Sep 16
Authors:
,
,

Abstract

Specialist training without explicit reasoning supervision implicitly selects latent reasoning trajectories that govern distilled students' specialization and generalization trade-offs.

Specialist distillation effectively transfers domain expertise to student models via teacher-generated reasoning trajectories. However, when these specialists are trained solely on question--answer pairs without explicit reasoning supervision, what governs the trajectories they generate? In this work, we show that specialist optimization implicitly selects from this latent trajectory space. To isolate and observe this latent distribution, we leverage student distillation not as a downstream goal, but as an agnostic probe---since students inherit no parameterization or optimization constraints from the specialist, inheriting only the sampled trajectories themselves. Through this probe, our empirical analysis unveils a tight governing relationship: across 27 specialist--student pairings, their specialization--generalization profiles correlate exceptionally strongly. Crucially, explicitly controlling the specialist's distributional drift systematically shifts both the teacher and its distilled student along a controllable trade-off between domain precision and general-capability retention. Across chemistry, physics, and multilingual settings, distilled students systematically reflect these specialist-induced profiles, even across divergent model families. Our findings establish a new view of specialist training: when gold reasoning is absent, tuning choices directly control the latent supervision passed to downstream models.

Community

A specialist is trained on answers only, with no reasoning supervision.

So what determines the reasoning trajectories it later generates to teach a student?

We find that specialist training itself implicitly selects these latent trajectories. Across 27 specialist–student pairs, students inherit their teachers’ specialization–generalization profiles—even across different model families.

And by controlling how the specialist is trained, we can systematically control what kind of reasoning supervision gets passed downstream.

Sign up or log in to comment

Get this paper in your agent:

hf papers read 2609.13770
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 0

No model linking this paper

Cite arxiv.org/abs/2609.13770 in a model README.md to link it from this page.

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2609.13770 in a dataset README.md to link it from this page.

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2609.13770 in a Space README.md to link it from this page.

Collections including this paper 0

No Collection including this paper

Add this paper to a collection to link it from this page.