jev-control-es: an extra-small, non-autoregressive decision model
A bidirectional-encoder decision model, the "extra small" sibling of
jev-control-core (0.8B decoder). Same target as that
model -- the decision sites inside a real agent harness (guardrail/injection gates, tool routing,
context ranking, answer sufficiency, triage, moderation, claim verification, next-action, entity
matching, escalate-to-human; see JevControl's Guide) -- but
a different, non-autoregressive architecture borrowed from
Laya (Convai, Apache 2.0): every option is scored at
its own dedicated [MASK] token in a single bidirectional pass and softmaxed over the question's
options, instead of a decoder's letter-slot readout. No LM head, no letter-token restriction, no fixed
26-option ceiling -- the option text is never generated, only scored in place.
Base: answerdotai/ModernBERT-base (149M,
bidirectional, 22 layers, 8,192-token context), fully fine-tuned, plus a linear scorer head trained from
scratch on the backbone's last hidden state at each [MASK] position. Laya's own head additionally runs
the option-marker positions through two extra from-scratch transformer layers before scoring; this
repository uses the simpler single-linear-layer head (a natural v2 refinement once these numbers have a
baseline) and ModernBERT-base rather than Laya's -large (395M), since "extra small" is the point
of this variant.
How it works
Context: <state>
Question: <question prompt>
Options:
- <option 1>: [MASK]
- <option 2>: [MASK]
...
The backbone runs once over this whole prompt; each [MASK] position's hidden state is projected to a
scalar by the linear head; the scalars for one question's options are softmaxed. Served through the same
/v1/systemone wire format as spark-s1 and jev-control-core (open_spark_jev.serve.gateway, which detects a
control_es.json marker file and routes to this architecture's scorer).
Training
Fully fine-tuned (backbone + head, 149M parameters) for 4 epochs on 5,000 rows (500/family) from
os_datagen.control's ten decision-site families, pinned to commit 92d92c9 -- not the same,
larger data snapshot jev-control-core (this family's 0.8B sibling) ships on; see History below for why
this model was deliberately rolled back to the smaller generation. lr 5e-5, batch 32, cosine schedule,
seed: 0 (deterministic -- configs/train/sft_control_jev_es.yaml), the same KL + Brier-regularised loss
as spark-s1's sft.py (architecture-agnostic: it only needs restricted per-option logits and a target
distribution, so it was reused unchanged). Every choice-type family's option list is rendered in a
per-row-random order -- see the history note below, this was the fix for a real bug, not a stylistic
choice.
Evaluation
Own held-out splits (scored at written option order and averaged over 3 random re-permutations):
| split | accuracy | accuracy (mean over 3 permutations) | ECE (calibrated) |
|---|---|---|---|
| test_locked | 1.000 | 1.000 | 0.000 |
| challenge | 1.000 | 1.000 | 0.000 |
Latency, isolated single-decision calls, one NVIDIA GB10: 7.7 ms per decision (median) -- faster than Laya itself (39.5 ms on a T4) and less than half jev-control-core's 18.5 ms, at less than a fifth the parameters.
Real harness (JevControl, support_desk demo, 203 tasks, against a Gemma-4-E4B baseline):
| arm | accuracy | verdict |
|---|---|---|
| Gemma 4 (prompted) | 0.892 | baseline |
| jev-control-es | 0.695 | worse (Δ -0.197), but 4x better than the first trained version |
Per site: injection 0.872, route 0.824, sufficiency 0.770, relevance 0.188 (see Limitations -- this is the one site that did not respond well to the two fixes below, and got worse, not better, when we tried a third fix -- see History).
History: two real bugs, not one vague "domain gap"
The first trained version of this model scored 0.167 on the real harness, with injection collapsed to always answering "false" and sufficiency worse than a coin flip -- outright majority-class collapse, not a noisy transfer attempt. Two specific, diagnosed causes accounted for almost all of it:
- Templated, fixed-vocabulary synthetic state in the
route/sufficient/relevancefamilies (abstract fact-name lists instead of realistic customer-message and article prose). Rewriting these three families to use varied greetings/sign-offs and a 12-topic realistic KB corpus was necessary but, for this architecture, not sufficient on its own -- injection and sufficiency were still badly broken after this fix alone (0.128 overall, injection specifically fell further, to 0.128 from 0.882). - Fixed option order in the four choice-type families (
tool_routing,moderation_class,next_action,escalate_human) let the model learn "the answer is at position N" instead of reading option text -- invisible on our own eval, which never varied the order, but catastrophic against a real caller's own option order. Randomizing option order per training row recovered injection to 0.887 and took the overall real-harness number from 0.128 to 0.685.
relevance (context ranking, a score question) is the one site that stayed weak through both fixes:
driven by a strong bias toward predicting "fully relevant" regardless of content (truth=0 but predicted
1 or 2 in 417 of 723 graded calls in the shipped run). See Limitations.
Two follow-up fixes were tried and rejected -- both made things worse, and both are worth knowing before you retrain this line further:
- Laya's own extra-transformer-layers scorer head (two from-scratch
nn.TransformerEncoderlayers across the option-marker positions before the final linear projection, matching Laya's published architecture more closely) was implemented and tried, hypothesising it would let the model compare options against each other and fixrelevance's bias. It instead collapsedinjectionfrom 0.887 to 0.246 and dropped the overall real-harness number to 0.172. Reverted. - Scaling training data 4x (20,000 rows / 2,000 per family, much wider template and topic pools --
the same change that took
jev-control-core's real-harness number from 0.813 to 0.837) was tried next, on the theory thatinjection's cross-retrain instability (0.887 in one run, as low as 0.128-0.25 in others, despite ~1.0 in-distribution accuracy every time) was a data-diversity problem, not an architecture one. On this architecture specifically, it made things worse:injectioncollapsed to 0.143 and the overall real-harness number to 0.138, despite the wider data and a clean, unchanged own-split accuracy (0.996). This checkpoint (jev-control-es, the one actually shipped in this repository) was rolled back to the original 500-rows/family generation (os_datagen.control, commit92d92c9) instead, which reproduces 0.695 deterministically (seed: 0inconfigs/train/sft_control_jev_es.yaml).
The net conclusion: this architecture's injection (noul) readout is unusually sensitive to training
run-to-run variance in a way jev-control-core's decoder readout is not, and neither "more transformer
layers" nor "more data" has fixed it so far -- see Limitations for what to try next if you pick this up.
Limitations
relevance(context ranking) is substantially below jev-control-core's 0.770 on the same real-harness site (0.188 here) and below the other three sites on this model (0.77-0.87). The failure mode is a strong bias toward the top score regardless of article content, not random noise. Score questions are a[MASK]-per-level readout the way Laya's own card also flags as its weakest primitive. Two attempted fixes (Laya's own extra transformer layers before scoring, and 4x more/wider training data) both made things worse rather than better on this architecture -- see History above. Worth trying next: a higherlambda_brierspecifically forscore-type questions (this site is the only one using that question type in training), or a dedicated ranking/contrastive loss across an article's candidate scores rather than the shared KL+Brier loss used for all three question types.injectionaccuracy has varied 0.13-0.89 across different retrains of this architecture despite ~1.0 in-distribution accuracy every time -- the shipped 0.872 number is one specific, reproducible checkpoint (seed: 0,os_datagen.controlcommit92d92c9, 500 rows/family), not a stable property of the architecture/data combination. If you retrain this model, re-run the real harness before trusting the number, even if the own-split eval looks perfect.- No reinforcement-learning stage. Laya's own card attributes part of its calibration edge to an RLCD stage (policy reports a distribution, reward is a strictly proper scoring rule, REINFORCE with a group-mean baseline) that this repository does not implement.
- vLLM does not support an arbitrary custom classification head on an encoder, so this model does not
serve through vLLM (neither does Laya, for the same reason) -- it serves through plain
transformers, which runs on CUDA, CPU and Apple Silicon (MPS) natively. - Per-question-type temperatures were fitted on the synthetic calibration split only; refit on your own labelled traffic before trusting confidence-gated escalation.
Files
config.json, tokenizer.json/tokenizer_config.json, model.safetensors (backbone), score_head.pt
(the linear scorer head's state dict), calibration.json, control_es.json (architecture marker read
by open_spark_jev.experimental.control_es.is_control_es).
Reproduction
datagen-pipeline/src/os_datagen/control/, configs/train/sft_control_jev_es.yaml,
open_spark_jev/train/sft_control_es.py, open_spark_jev/experimental/control_es.py in
abhishek085/open-spark-jev.
To reproduce this exact checkpoint: check out commit 92d92c9 of
datagen-pipeline/src/os_datagen/control/families.py specifically (the repository's main branch has
since moved to a wider/larger generation for jev-control-core; see History above for why this model
was deliberately kept on the earlier one), generate with
python -m os_datagen.control --out data/synthetic/control_v1 --n-train 500 --n-eval 150, then
python -m open_spark_jev.train.sft_control_es --config configs/train/sft_control_jev_es.yaml
(seed: 0 makes this deterministic).
- Downloads last month
- 10
Model tree for abhishek085/jev-control-es
Base model
answerdotai/ModernBERT-base