jev-control-es: an extra-small, non-autoregressive decision model

A bidirectional-encoder decision model, the "extra small" sibling of jev-control-core (0.8B decoder). Same target as that model -- the decision sites inside a real agent harness (guardrail/injection gates, tool routing, context ranking, answer sufficiency, triage, moderation, claim verification, next-action, entity matching, escalate-to-human; see JevControl's Guide) -- but a different, non-autoregressive architecture borrowed from Laya (Convai, Apache 2.0): every option is scored at its own dedicated [MASK] token in a single bidirectional pass and softmaxed over the question's options, instead of a decoder's letter-slot readout. No LM head, no letter-token restriction, no fixed 26-option ceiling -- the option text is never generated, only scored in place.

Base: answerdotai/ModernBERT-base (149M, bidirectional, 22 layers, 8,192-token context), fully fine-tuned, plus a linear scorer head trained from scratch on the backbone's last hidden state at each [MASK] position. Laya's own head additionally runs the option-marker positions through two extra from-scratch transformer layers before scoring; this repository uses the simpler single-linear-layer head (a natural v2 refinement once these numbers have a baseline) and ModernBERT-base rather than Laya's -large (395M), since "extra small" is the point of this variant.

How it works

Context: <state>
Question: <question prompt>
Options:
- <option 1>: [MASK]
- <option 2>: [MASK]
...

The backbone runs once over this whole prompt; each [MASK] position's hidden state is projected to a scalar by the linear head; the scalars for one question's options are softmaxed. Served through the same /v1/systemone wire format as spark-s1 and jev-control-core (open_spark_jev.serve.gateway, which detects a control_es.json marker file and routes to this architecture's scorer).

Training

Fully fine-tuned (backbone + head, 149M parameters) for 4 epochs on 5,000 rows (500/family) from os_datagen.control's ten decision-site families, pinned to commit 92d92c9 -- not the same, larger data snapshot jev-control-core (this family's 0.8B sibling) ships on; see History below for why this model was deliberately rolled back to the smaller generation. lr 5e-5, batch 32, cosine schedule, seed: 0 (deterministic -- configs/train/sft_control_jev_es.yaml), the same KL + Brier-regularised loss as spark-s1's sft.py (architecture-agnostic: it only needs restricted per-option logits and a target distribution, so it was reused unchanged). Every choice-type family's option list is rendered in a per-row-random order -- see the history note below, this was the fix for a real bug, not a stylistic choice.

Evaluation

Own held-out splits (scored at written option order and averaged over 3 random re-permutations):

split accuracy accuracy (mean over 3 permutations) ECE (calibrated)
test_locked 1.000 1.000 0.000
challenge 1.000 1.000 0.000

Latency, isolated single-decision calls, one NVIDIA GB10: 7.7 ms per decision (median) -- faster than Laya itself (39.5 ms on a T4) and less than half jev-control-core's 18.5 ms, at less than a fifth the parameters.

Real harness (JevControl, support_desk demo, 203 tasks, against a Gemma-4-E4B baseline):

arm accuracy verdict
Gemma 4 (prompted) 0.892 baseline
jev-control-es 0.695 worse (Δ -0.197), but 4x better than the first trained version

Per site: injection 0.872, route 0.824, sufficiency 0.770, relevance 0.188 (see Limitations -- this is the one site that did not respond well to the two fixes below, and got worse, not better, when we tried a third fix -- see History).

History: two real bugs, not one vague "domain gap"

The first trained version of this model scored 0.167 on the real harness, with injection collapsed to always answering "false" and sufficiency worse than a coin flip -- outright majority-class collapse, not a noisy transfer attempt. Two specific, diagnosed causes accounted for almost all of it:

  1. Templated, fixed-vocabulary synthetic state in the route/sufficient/relevance families (abstract fact-name lists instead of realistic customer-message and article prose). Rewriting these three families to use varied greetings/sign-offs and a 12-topic realistic KB corpus was necessary but, for this architecture, not sufficient on its own -- injection and sufficiency were still badly broken after this fix alone (0.128 overall, injection specifically fell further, to 0.128 from 0.882).
  2. Fixed option order in the four choice-type families (tool_routing, moderation_class, next_action, escalate_human) let the model learn "the answer is at position N" instead of reading option text -- invisible on our own eval, which never varied the order, but catastrophic against a real caller's own option order. Randomizing option order per training row recovered injection to 0.887 and took the overall real-harness number from 0.128 to 0.685.

relevance (context ranking, a score question) is the one site that stayed weak through both fixes: driven by a strong bias toward predicting "fully relevant" regardless of content (truth=0 but predicted 1 or 2 in 417 of 723 graded calls in the shipped run). See Limitations.

Two follow-up fixes were tried and rejected -- both made things worse, and both are worth knowing before you retrain this line further:

  • Laya's own extra-transformer-layers scorer head (two from-scratch nn.TransformerEncoder layers across the option-marker positions before the final linear projection, matching Laya's published architecture more closely) was implemented and tried, hypothesising it would let the model compare options against each other and fix relevance's bias. It instead collapsed injection from 0.887 to 0.246 and dropped the overall real-harness number to 0.172. Reverted.
  • Scaling training data 4x (20,000 rows / 2,000 per family, much wider template and topic pools -- the same change that took jev-control-core's real-harness number from 0.813 to 0.837) was tried next, on the theory that injection's cross-retrain instability (0.887 in one run, as low as 0.128-0.25 in others, despite ~1.0 in-distribution accuracy every time) was a data-diversity problem, not an architecture one. On this architecture specifically, it made things worse: injection collapsed to 0.143 and the overall real-harness number to 0.138, despite the wider data and a clean, unchanged own-split accuracy (0.996). This checkpoint (jev-control-es, the one actually shipped in this repository) was rolled back to the original 500-rows/family generation (os_datagen.control, commit 92d92c9) instead, which reproduces 0.695 deterministically (seed: 0 in configs/train/sft_control_jev_es.yaml).

The net conclusion: this architecture's injection (noul) readout is unusually sensitive to training run-to-run variance in a way jev-control-core's decoder readout is not, and neither "more transformer layers" nor "more data" has fixed it so far -- see Limitations for what to try next if you pick this up.

Limitations

  • relevance (context ranking) is substantially below jev-control-core's 0.770 on the same real-harness site (0.188 here) and below the other three sites on this model (0.77-0.87). The failure mode is a strong bias toward the top score regardless of article content, not random noise. Score questions are a [MASK]-per-level readout the way Laya's own card also flags as its weakest primitive. Two attempted fixes (Laya's own extra transformer layers before scoring, and 4x more/wider training data) both made things worse rather than better on this architecture -- see History above. Worth trying next: a higher lambda_brier specifically for score-type questions (this site is the only one using that question type in training), or a dedicated ranking/contrastive loss across an article's candidate scores rather than the shared KL+Brier loss used for all three question types.
  • injection accuracy has varied 0.13-0.89 across different retrains of this architecture despite ~1.0 in-distribution accuracy every time -- the shipped 0.872 number is one specific, reproducible checkpoint (seed: 0, os_datagen.control commit 92d92c9, 500 rows/family), not a stable property of the architecture/data combination. If you retrain this model, re-run the real harness before trusting the number, even if the own-split eval looks perfect.
  • No reinforcement-learning stage. Laya's own card attributes part of its calibration edge to an RLCD stage (policy reports a distribution, reward is a strictly proper scoring rule, REINFORCE with a group-mean baseline) that this repository does not implement.
  • vLLM does not support an arbitrary custom classification head on an encoder, so this model does not serve through vLLM (neither does Laya, for the same reason) -- it serves through plain transformers, which runs on CUDA, CPU and Apple Silicon (MPS) natively.
  • Per-question-type temperatures were fitted on the synthetic calibration split only; refit on your own labelled traffic before trusting confidence-gated escalation.

Files

config.json, tokenizer.json/tokenizer_config.json, model.safetensors (backbone), score_head.pt (the linear scorer head's state dict), calibration.json, control_es.json (architecture marker read by open_spark_jev.experimental.control_es.is_control_es).

Reproduction

datagen-pipeline/src/os_datagen/control/, configs/train/sft_control_jev_es.yaml, open_spark_jev/train/sft_control_es.py, open_spark_jev/experimental/control_es.py in abhishek085/open-spark-jev.

To reproduce this exact checkpoint: check out commit 92d92c9 of datagen-pipeline/src/os_datagen/control/families.py specifically (the repository's main branch has since moved to a wider/larger generation for jev-control-core; see History above for why this model was deliberately kept on the earlier one), generate with python -m os_datagen.control --out data/synthetic/control_v1 --n-train 500 --n-eval 150, then python -m open_spark_jev.train.sft_control_es --config configs/train/sft_control_jev_es.yaml (seed: 0 makes this deterministic).

Downloads last month
10
Safetensors
Model size
0.1B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for abhishek085/jev-control-es

Finetuned
(1502)
this model