Laya Stop-Completion Judge
Fine-tune of convaiinnovations/laya (ModernBERT-large 421M encoder + typed-decision head) specialized on a single judgment: did an AI coding agent's turn actually finish, or did it stop with announced work still undelivered?
Built to guard the Stop hook of a Claude Code agent fleet: when an agent tries to end its turn, this model scores P(incomplete) on (user request, final assistant message) in ~25 ms on Apple silicon. Turns scoring high are blocked and sent back to either finish the work or explicitly hand back.
Why this exists
Some models habitually end their turn while still announcing pending work: "Fixed the locator. Next I'll run the testsโฆ" โ and then just stop. This fleet routes most routine work to open-weights models (Kimi K3, GLM) through an internal proxy, and K3 in particular did this constantly: turn after turn ending on a promise, with the actual work never started. Frontier models rarely exhibit it; cheaper open-weights models do it enough that unattended background sessions can't be trusted without a guard.
The first fix was a regex (stop_guard.py, 2026-08): match announcement
phrases in the final sentence and block. It works but only knows its own
vocabulary โ it misses danglers phrased outside it ("Checking for stray
em-dashes before calling it done:", "Not in the first page โ checking the
next:") and can't weigh context (a message reporting queued work is not the
same as one announcing personal work).
This checkpoint replaces pattern-matching with judgment: it reads what was asked against what was delivered and scores P(incomplete). The practical payoff is routing freedom โ send routine work to the cheapest model that can do it, and let a 25 ms local judge compensate for the premature-stop habit instead of paying frontier prices just to get reliable turn endings.
Results (held-out real turns, n=315, 59 positive)
| model | AUROC | mean P(inc) on complete | mean P(inc) on incomplete |
|---|---|---|---|
| base laya, legacy prompt | 0.64 | 0.29 | 0.35 |
| base laya, packed state | 0.57 | 0.31 | 0.33 |
| this checkpoint | 0.85 | 0.08 | 0.49 |
The base checkpoint is chance-level at auditing agentic turns โ the fine-tune
is the whole signal. ECE 0.057, accuracy@0.5 0.87. Per-epoch metrics in
metrics.json.
Training data
2,627 real turn-ending pairs (user request, final assistant message) harvested from Claude Code session transcripts (September 2026, one operator's fleet of mostly-autonomous background sessions). A turn-end is the last text-bearing assistant message before the next real user message โ not mid-turn narration (getting this wrong in the mining pass produced an 78%-inverted label distribution; see Limitations).
Labels were distilled from Kimi K3 judging completion against a fixed rubric (announced-but-unreported next steps = incomplete; reported results, direct answers, explicit handbacks = complete), with manual overrides on hook-intercepted turns; labels below 0.5 judge confidence were dropped. The transcripts are private and are not released.
Training: encoder frozen, decision head only (26.5M params), strictly-proper scoring loss (log + spherical, the base repo's own objective), option-order shuffling, 8 epochs, bf16 on Apple MPS, incomplete class oversampled 2:1.
Usage
import laya
from state_pack import pack_state, COMPLETION_QUESTION # shipped in this repo
judge = laya.load("tampajohn/laya-stop-completion-judge", device="mps") # or cuda / cpu
res = judge.predict(
pack_state(user_text, final_assistant_text),
{"completion": COMPLETION_QUESTION},
)
ans = res["answers"]["completion"]
p_incomplete = ans["probabilities"]["incomplete"]
Use the shipped state_pack.py. The model was trained on this exact state
packing (final-message tail first) and question text. The base SDK's default
serialization truncates the state head-first, so long final messages lose the
tail โ which is where "next I'llโฆ" announcements live. That mismatch is part
of why the base model fails this task.
Deployment notes (as used in production)
- Block when
p_incomplete >= 0.75; calibrate on your own traffic before blocking. - Suppress blocks when the final message's last line is a direct question โ clarification questions are explicit handbacks and this model over-fires on them (observed 0.68โ0.85). Shape checks belong in code, not the model.
- Feed it the turn's true final message from the agent harness's stop payload, not from a transcript that may lag โ a mid-turn status line fed in by mistake scored 0.85 and produced a false block.
- Runs happily beside the base model in one daemon (~1.7 GB weights, ~25 ms predict on M-series MPS).
Limitations
- Single-operator distribution: autonomous background agent sessions with
strong house conventions (
result:completion lines, scheduled monitoring loops). Expect domain shift on interactive chat or other agent harnesses. - LLM-distilled labels: biases of the labeling judge are inherited; ~10% of kept labels were borderline.
- English only; sees โค1024 tokens of packed state.
- Known failure modes: clarification-question handbacks (guard above), and stale-input inflation (feed it the real final message).
- Not a general-purpose router โ this checkpoint is trained for one question. For general typed decisions, use the base model.
Provenance
Built 2026-09-25/26 as the completion judge for a layad hook daemon guarding
a Claude Code fleet (stop / subagent-stop / notification hooks as pure curl
one-liners to a local preloaded server). The full pipeline โ transcript turn
mining, LLM labeling, head fine-tune, shadow deployment, live blocking โ was
developed against real production traffic; the base model showed no measurable
separation on this task (0.15 vs 0.15 P(incomplete) on synthetic pairs,
AUROC โ 0.6 on real held-out turns), motivating the fine-tune.
Model tree for tampajohn/laya-stop-completion-judge
Base model
convaiinnovations/laya