lod-stor-4B

A 4B option-scoring model. You send a state and one or more typed questions, each with its own options and criteria. For every question it returns a calibrated probability over your options, plus a separate confidence that the answer is among them at all. No generated tokens; runs on a CPU. One of two sizes: lod-lille-0.6B (Qwen3-0.6B-Base) and lod-stor-4B (Qwen3-4B-Base), same interface and code.

Code, training recipes and the corpus pipeline: github.com/mrn-dk/lod.

  • Independent questions: each question attends to the state and to itself, never to another question, so adding or removing a question never changes the rest.
  • Independent options: each option attends to the state, its question and itself, from the same start position. No option's score depends on which other options are listed or in what order.
  • Options are compared once: a small comparison layer looks at the top 64 options together and corrects their scores; the rest keep their own.
  • Only your options: the softmax covers your options and nothing else, at temperature 1. The training loss is a proper scoring rule, and a temperature fitted afterwards made held-out calibration worse, so none is applied.
  • Criteria are read: option keys are arbitrary identifiers; the criteria text carries the meaning.
  • Proper-scoring training: cross-entropy against observed outcome frequencies (soft targets where they exist). No RL.

Quick start

pip install torch transformers
from transformers import AutoModel

model = AutoModel.from_pretrained("mrn-dk/lod-stor-4B", trust_remote_code=True).eval()

response = model.score(
    state={"channel": "email",
           "subject": "Deploy failing since 4.2.1",
           "body": "Since upgrading to 4.2.1 our nightly deploy fails at the migration "
                   "step with 'relation users_pkey already exists'. Rolling back to "
                   "4.2.0 fixes it. We have 200 seats and this blocks our release on "
                   "Friday.",
           "plan": "enterprise"},
    questions={
        "kind": {
            "type": "choice",
            "instructions": "What kind of request is this?",
            "criteria": {"bug":      "Reports something that is broken",
                         "question": "Asks how to do something",
                         "feature":  "Requests new functionality",
                         "billing":  "About payment or invoices"}},
        "needs_engineer": {
            "type": "noul",
            "instructions": "Does resolving this require someone who can read the codebase?",
            "criteria": {"true":  "Requires engineering investigation",
                         "false": "A support agent can resolve it from documentation"}},
    },
)

kind = response["answers"]["kind"]
print({k: round(v, 3) for k, v in kind["probabilities"].items()}, round(kind["confidence"], 3))
# {'bug': 0.966, 'question': 0.027, 'feature': 0.004, 'billing': 0.003} 0.92
print(round(response["answers"]["needs_engineer"]["noul"], 3))
# 0.687
print(response["usage"])
# {'input_tokens': 173, 'output_tokens': 0, 'state_tokens': 91, 'state_truncated': False}

confidence is the head's estimate that the right answer is among the listed options; noul is the probability that the statement is true.

Interface

question type asks returns
choice which option is correct choice, probabilities, confidence
noul whether a statement is true noul (a probability)
score a position on a 2–10 point scale score, probabilities, confidence
  • The model never sees the question id; put the full question in instructions.
  • No option cap. When the options do not fit beside the state, each question is scored in option shards against the state encoded once. Because options are independent, that gives the same answer as one pass.
  • Up to 32,768 tokens accepted (the backbone's window): the state is read up to 30,720 tokens, and the state plus the questions of one pass up to 32,768.
  • Tested up to 30,000-token states (the long-context check above); 256 options in one question (synthetic) and 988 test questions with 129+ options. Trained on states up to 30,000 tokens and 32,768 in total, mostly short.
  • The state is never cut silently. A state over the budget raises ValidationError (field == "state"). Pass state_overflow="truncate" to score the part that fits; usage.state_truncated and usage.state_tokens then say how much was read.
  • usage also reports input_tokens; output_tokens is always 0.
  • fp32 by default; dtype=torch.bfloat16 halves memory and moves the example slightly (probabilities unchanged to 3 decimals; noul 0.692). model.to("cuda") for GPU.
  • config.json sets temperature: 1.0 (no fitted temperature) and confidence_mode: head; score() applies both.

Evaluation

139,543 test questions over 447 tasks. The test tasks' families, and the development families used to choose the checkpoint, never appear in training. 95 % bootstrap intervals are in brackets.

accuracy NLL ECE
all questions 0.789 [0.786, 0.791] 0.665 0.036 [0.034, 0.038]
mean over tasks 0.685 [0.661, 0.710] 0.931 —
labels from people or real systems (70,997 questions, 375 tasks) 0.738 0.798 0.036
labels computed by code (68,546 questions, 72 tasks) 0.842 0.539 0.038
  • Selective prediction: when the top option has p ≥ 0.9 (47.5 % of questions) it is wrong 5.5 % of the time. The confidence head ranks answers better than the top probability (AURC 0.045 vs 0.069 on held-out dev tasks).
  • Option order: reversing and shuffling the options of 41 questions moves a probability by 0.0007 on average and at most 0.010 (bf16 GPU arithmetic), and never changes the answer; in fp32 the scores are order-independent by construction.
  • Many options: accuracy on test questions with 129 or more options is 0.475 (988 questions). Synthetic look-up, 10 questions per option count, accuracy by N — 2: 1.0, 4: 1.0, 8: 1.0, 16: 1.0, 32: 1.0, 64: 1.0, 128: 1.0, 256: 1.0.
  • Long states: controlled look-up at fixed state lengths, 180 questions each (chance 0.306), accuracy — 1k: 0.85, 2k: 0.82, 4k: 0.81, 8k: 0.82, 16k: 0.79, 30k: 0.81.
  • Abstention: with the needed evidence withheld, mean confidence drops from 0.84 to 0.52; for claims about entities never seen in training it is 0.97 when the state supports the claim and 0.30 when the state does not mention it.

Limitations

  • Accuracy varies a lot by domain. The weakest on the test tasks: reasoning and puzzles (0.50), document structure (0.57), security advisory triage (0.59), text classification and taxonomy (0.61).
  • Probabilities are calibrated on average (ECE 0.036), but per task they can be far off (mean per-task ECE 0.19). On questions with a known true probability the model tracks it (r 0.78, mean absolute error 0.089) but not closely enough to read as an exact posterior.
  • It cannot answer off-list. If the right answer is missing, only confidence is meaningful.
  • Long states are a tail of the training mix; beyond 30,000 tokens the numbers above are unmeasured.
  • Only the top 64 options by first-stage score are compared with each other; the others are scored on their own.
  • English only.

Training

Qwen/Qwen3-4B-Base with LoRA (r 64, alpha 128, dropout 0.05), merged. 8,000 steps at learning rate 1e-4, batches of up to 49,152 tokens; checkpoint chosen by out-of-distribution dev NLL. No temperature is fitted: one fitted on the held-out dev families (with an option-count slope) lowered NLL slightly but made calibration worse. The confidence head is trained afterwards against the frozen scoring model and reads the shortlisted options' hidden states plus the shape of the distribution.

Data: 793,920 questions over 2,750 tasks from 51 sources in 25 domains: text classification, rule application, affect, reasoning, forecasting, grounding, security, software, NLI, law, tables, documents, logs, browser and sensor states, tool calls, entity resolution, code diffs, scheduling, multi-hop ranking, multilingual text, exact posteriors, many-option linking, broad real-language judgements and long documents. Labels come from people or real systems where they exist and are computed by code for the generated rows. The Decision Index 0.2 benchmark is excluded (its datasets by name, and every state checked against its items).

Citation

@software{lod_2026,
  title  = {lod-stor-4B: a prefill-only option-scoring model for calibrated decisions},
  year   = {2026},
  note   = {Qwen/Qwen3-4B-Base + LoRA, independent questions and options, a comparison
            layer over the shortlist, proper-scoring fine-tuning, separate confidence
            head}
}
Downloads last month
8
Safetensors
Model size
4B params
Tensor type
BF16
·
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for mrn-dk/lod-stor-4B

Finetuned
(481)
this model

Evaluation results

  • Accuracy (micro) on Lod held-out test tasks
    self-reported
    0.789
  • NLL vs soft targets (micro) on Lod held-out test tasks
    self-reported
    0.665
  • Expected calibration error on Lod held-out test tasks
    self-reported
    0.036
  • AURC, confidence head on Lod held-out test tasks
    self-reported
    0.045