Sev — typed-decisions, CE-only, 1024/256
A ModernBERT-large decision model fine-tuned on LocalLLaMA/typed-decisions with plain cross-entropy against the teacher's soft targets, at the documented 1024-token context / 256-token option budget.
It is the best-performing checkpoint from the Sev study (Dissecting RLCD), where I took apart the training method behind TypeSafe's Jev and its open reproduction, Laya. The short version of that study: Laya's RL term is a noise-smoothed cross-entropy gradient, and on this benchmark plain CE matches or beats it on every proper score.
Test-set results (2,000 decisions, 400 cases):
| Metric | Value |
|---|---|
| Accuracy (vs hard label) | 0.7885 |
| Brier (vs soft targets) | 0.0495 |
| NLL (vs soft targets) | 0.8581 |
| Soft accuracy | 0.5236 |
| Mean confidence | 0.6444 |
| Mean target max | 0.6589 |
| Score MAE (ordinal questions) | 0.2115 |
| Fitted temperature (choice / score / noul) | 1.116 / 1.070 / 1.120 |
For reference, the laya-typed-decisions checkpoint reports 0.766 accuracy and 0.062 Brier; TypeSafe's published Jev 1.13.0 figure is 0.727.
Model description
A bidirectional encoder plus a small transformer head that reads one logit per masked option marker:
[CLS] <type> question: <instr> [SEP] [MASK] opt0 [MASK] opt1 ... [SEP] <state> [SEP]
|
ModernBERT-large encoder (bidirectional)
|
h += type_emb(qtype)
|
2 x TransformerEncoderLayer (pre-norm)
|
gather h at [MASK] positions -> scorer MLP -> 1 logit/option
One input row per typed question. Question types are choice (pick one of K named options), score (ordinal), and noul (binary yes/no). Option count is per-row, not fixed.
- Encoder:
answerdotai/ModernBERT-large(395M) - Head: 2 pre-norm transformer layers, dropout 0.1
- Scorer: LayerNorm, Linear, GELU, Linear(->1)
- Total: ~421M parameters
- Sequence budget: 1024 tokens total, 256 for the question + option block
The output is a probability distribution over the row's options. For the score type, the expected level is sum i * p_i.
Intended use
Structured, closed-set probabilistic decisions over short English text: routing, triage, classification with calibrated confidence, and ordinal rating tasks where you control the option set.
Out of scope: open-ended generation (this model does not generate text), option sets larger than ~255, languages other than English (the base is English; a multilingual Laya variant exists), and any use where the option wording and the state distribution differ sharply from the training workflows.
How to use
This is a custom architecture, not a transformers AutoModel. Load it with the code from the reproduction repo:
git clone https://github.com/LakoreAI/sev
cd rlcd-reverse-engineering && uv sync
from pathlib import Path
import torch
from huggingface_hub import snapshot_download
import sys
sys.path.insert(0, "rlcd-reverse-engineering") # or pip install -e it
from src.pipelines.infer import load_model, infer
d = snapshot_download("LakoreAI/sev")
model, cfg, tokenizer = load_model(Path(d) / "model.safetensors", torch.device("cuda"))
result = infer(
Path(d) / "model.safetensors",
state='{"task": "triage", "ticket": "customer cannot log in after password reset"}',
question='{"type": "choice", "instructions": "Which queue should this go to?",'
' "criteria": {"billing": "payment or invoice", "auth": "login or account access",'
' "network": "connectivity"}}',
)
print(result["predicted_key"], result["probabilities"])
load_model reads model_config.json to rebuild the architecture and temperatures.json to apply the fitted per-(type, K-bucket) temperature at inference, so the reported probabilities are already calibration-scaled.
Training data
LocalLLaMA/typed-decisions (Apache-2.0): 1,200 train cases and 400 test cases across four workflows (agent-trace observability, customer service, invoice processing, security incidents), flattened to ~5,400 training decisions. Targets are the dataset's soft teacher distributions, not the hard labels.
Training procedure
- Initialised from
convaiinnovations/laya. - Loss: soft cross-entropy against the teacher distributions only (no RL term).
- 4 epochs, effective batch 64 (micro-batch 8 x 4 accumulation on a single 24 GB GPU), AdamW, lr 2.5e-5 encoder / 1e-4 head, weight decay 0.01, cosine decay to 1e-6, gradient clipping 1.0, bf16.
- No early stopping; the final-epoch weights are evaluated.
- Post-hoc temperature fitted per (question type, option-count bucket) by NLL on a held-out 10% slice of the training cases.
Evaluation
Full test-set breakdown by question type (2,000 decisions):
| Type | n | Accuracy | Brier | NLL | Mean conf. | Mean target max |
|---|---|---|---|---|---|---|
| choice | 600 | 0.7550 | 0.0611 | 0.9818 | 0.6194 | 0.6345 |
| score | 800 | 0.7588 | 0.0511 | 1.0331 | 0.5817 | 0.5929 |
| noul | 600 | 0.8617 | 0.0356 | 0.5009 | 0.7532 | 0.7714 |
| overall | 2000 | 0.7885 | 0.0495 | 0.8581 | 0.6444 | 0.6589 |
A calibration caveat that matters here. On this benchmark the gold hard label is the teacher's argmax on 98.4% of rows, so hard-label ECE is not a valid calibration measure: a model that reproduces the teacher perfectly still scores 0.326 hard-label ECE, and over-sharpening lowers it. The raw hard-label ECE of this checkpoint is 0.144 (0.162 after temperature). Judge calibration against the soft targets instead, with Brier and NLL above. The fitted temperatures (all > 1) are the model's own signal that the raw logits were slightly over-sharp before scaling.
Relation to RLCD
This checkpoint is deliberately RL-free. In the accompanying study, adding Laya's RL term (a score-function estimate of a noise-smoothed proper score) did not improve accuracy and made the fitted temperature rise with the training noise scale. CE-only was at least as good on every proper score. See the linked repo for the full ablation, the estimator derivation, and the per-run result files.
Limitations
- Single benchmark; the four workflows are synthetic and teacher-generated.
- English only.
- Option sets above ~20 options degrade, because options share a fixed token budget.
- Accuracy differences of ~1 point are within seed noise on a dataset this size; three seeds gave 0.782 +/- 0.004 for the 512-token variant.
- The base checkpoint was already fine-tuned by the Laya authors; this model is a further fine-tune on the same benchmark's train split.
Links
- Code, configs, and per-run results: LakoreAI/sev
- All ablation checkpoints and metrics: minhleduc/rlcd-e2-checkpoints
- Base model: convaiinnovations/laya
- Dataset: LocalLLaMA/typed-decisions
Citation
@misc{leduc2026dissectingrlcd,
title = {Dissecting RLCD: What Reinforcement Learning Does (and Doesn't) Do for Calibrated Typed Decisions},
author = {Le Duc Minh},
year = {2026},
url = {https://github.com/LakoreAI/sev}
}
- Downloads last month
- 12
Model tree for LakoreAI/sev
Base model
convaiinnovations/layaDataset used to train LakoreAI/sev
Evaluation results
- Accuracy on LocalLLaMA/typed-decisionstest set self-reported0.788
- Brier (vs soft targets) on LocalLLaMA/typed-decisionstest set self-reported0.050
- NLL (vs soft targets) on LocalLLaMA/typed-decisionstest set self-reported0.858