Instructions to use mrn-dk/lod-stor-4B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use mrn-dk/lod-stor-4B with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("zero-shot-classification", model="mrn-dk/lod-stor-4B", trust_remote_code=True)# Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("mrn-dk/lod-stor-4B", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
lod-stor-4B
A 4B option-scoring model. You send a state and one or more typed questions, each
with its own options and criteria. For every question it returns a calibrated probability
over your options, plus a separate confidence that the answer is among them at all. No
generated tokens; runs on a CPU. One of two sizes: lod-lille-0.6B (Qwen3-0.6B-Base)
and lod-stor-4B (Qwen3-4B-Base), same interface and code.
Code, training recipes and the corpus pipeline: github.com/mrn-dk/lod.
- Independent questions: each question attends to the state and to itself, never to another question, so adding or removing a question never changes the rest.
- Independent options: each option attends to the state, its question and itself, from the same start position. No option's score depends on which other options are listed or in what order.
- Options are compared once: a small comparison layer looks at the top 64 options together and corrects their scores; the rest keep their own.
- Only your options: the softmax covers your options and nothing else, at temperature 1. The training loss is a proper scoring rule, and a temperature fitted afterwards made held-out calibration worse, so none is applied.
- Criteria are read: option keys are arbitrary identifiers; the criteria text carries the meaning.
- Proper-scoring training: cross-entropy against observed outcome frequencies (soft targets where they exist). No RL.
Quick start
pip install torch transformers
from transformers import AutoModel
model = AutoModel.from_pretrained("mrn-dk/lod-stor-4B", trust_remote_code=True).eval()
response = model.score(
state={"channel": "email",
"subject": "Deploy failing since 4.2.1",
"body": "Since upgrading to 4.2.1 our nightly deploy fails at the migration "
"step with 'relation users_pkey already exists'. Rolling back to "
"4.2.0 fixes it. We have 200 seats and this blocks our release on "
"Friday.",
"plan": "enterprise"},
questions={
"kind": {
"type": "choice",
"instructions": "What kind of request is this?",
"criteria": {"bug": "Reports something that is broken",
"question": "Asks how to do something",
"feature": "Requests new functionality",
"billing": "About payment or invoices"}},
"needs_engineer": {
"type": "noul",
"instructions": "Does resolving this require someone who can read the codebase?",
"criteria": {"true": "Requires engineering investigation",
"false": "A support agent can resolve it from documentation"}},
},
)
kind = response["answers"]["kind"]
print({k: round(v, 3) for k, v in kind["probabilities"].items()}, round(kind["confidence"], 3))
# {'bug': 0.966, 'question': 0.027, 'feature': 0.004, 'billing': 0.003} 0.92
print(round(response["answers"]["needs_engineer"]["noul"], 3))
# 0.687
print(response["usage"])
# {'input_tokens': 173, 'output_tokens': 0, 'state_tokens': 91, 'state_truncated': False}
confidence is the head's estimate that the right answer is among the listed options; noul is the probability that the statement is true.
Interface
| question type | asks | returns |
|---|---|---|
choice |
which option is correct | choice, probabilities, confidence |
noul |
whether a statement is true | noul (a probability) |
score |
a position on a 2–10 point scale | score, probabilities, confidence |
- The model never sees the question id; put the full question in
instructions. - No option cap. When the options do not fit beside the state, each question is scored in option shards against the state encoded once. Because options are independent, that gives the same answer as one pass.
- Up to 32,768 tokens accepted (the backbone's window): the state is read up to 30,720 tokens, and the state plus the questions of one pass up to 32,768.
- Tested up to 30,000-token states (the long-context check above); 256 options in one question (synthetic) and 988 test questions with 129+ options. Trained on states up to 30,000 tokens and 32,768 in total, mostly short.
- The state is never cut silently. A state over the budget raises
ValidationError(field == "state"). Passstate_overflow="truncate"to score the part that fits;usage.state_truncatedandusage.state_tokensthen say how much was read. usagealso reportsinput_tokens;output_tokensis always 0.- fp32 by default;
dtype=torch.bfloat16halves memory and moves the example slightly (probabilities unchanged to 3 decimals; noul 0.692).model.to("cuda")for GPU. config.jsonsetstemperature: 1.0(no fitted temperature) andconfidence_mode: head;score()applies both.
Evaluation
139,543 test questions over 447 tasks. The test tasks' families, and the development families used to choose the checkpoint, never appear in training. 95 % bootstrap intervals are in brackets.
| accuracy | NLL | ECE | |
|---|---|---|---|
| all questions | 0.789 [0.786, 0.791] | 0.665 | 0.036 [0.034, 0.038] |
| mean over tasks | 0.685 [0.661, 0.710] | 0.931 | — |
| labels from people or real systems (70,997 questions, 375 tasks) | 0.738 | 0.798 | 0.036 |
| labels computed by code (68,546 questions, 72 tasks) | 0.842 | 0.539 | 0.038 |
- Selective prediction: when the top option has p ≥ 0.9 (47.5 % of questions) it is wrong 5.5 % of the time. The
confidencehead ranks answers better than the top probability (AURC 0.045 vs 0.069 on held-out dev tasks). - Option order: reversing and shuffling the options of 41 questions moves a probability by 0.0007 on average and at most 0.010 (bf16 GPU arithmetic), and never changes the answer; in fp32 the scores are order-independent by construction.
- Many options: accuracy on test questions with 129 or more options is 0.475 (988 questions). Synthetic look-up, 10 questions per option count, accuracy by N — 2: 1.0, 4: 1.0, 8: 1.0, 16: 1.0, 32: 1.0, 64: 1.0, 128: 1.0, 256: 1.0.
- Long states: controlled look-up at fixed state lengths, 180 questions each (chance 0.306), accuracy — 1k: 0.85, 2k: 0.82, 4k: 0.81, 8k: 0.82, 16k: 0.79, 30k: 0.81.
- Abstention: with the needed evidence withheld, mean confidence drops from 0.84 to 0.52; for claims about entities never seen in training it is 0.97 when the state supports the claim and 0.30 when the state does not mention it.
Limitations
- Accuracy varies a lot by domain. The weakest on the test tasks: reasoning and puzzles (0.50), document structure (0.57), security advisory triage (0.59), text classification and taxonomy (0.61).
- Probabilities are calibrated on average (ECE 0.036), but per task they can be far off (mean per-task ECE 0.19). On questions with a known true probability the model tracks it (r 0.78, mean absolute error 0.089) but not closely enough to read as an exact posterior.
- It cannot answer off-list. If the right answer is missing, only
confidenceis meaningful. - Long states are a tail of the training mix; beyond 30,000 tokens the numbers above are unmeasured.
- Only the top 64 options by first-stage score are compared with each other; the others are scored on their own.
- English only.
Training
Qwen/Qwen3-4B-Base with LoRA (r 64, alpha 128, dropout 0.05), merged. 8,000 steps at learning rate 1e-4, batches of up to 49,152 tokens; checkpoint chosen by
out-of-distribution dev NLL. No temperature is fitted: one fitted on the held-out dev
families (with an option-count slope) lowered NLL slightly but made calibration worse.
The confidence head is trained afterwards against the frozen scoring model and reads the shortlisted options' hidden states plus the shape of the
distribution.
Data: 793,920 questions over 2,750 tasks from 51 sources in 25 domains: text classification, rule application, affect, reasoning, forecasting, grounding, security, software, NLI, law, tables, documents, logs, browser and sensor states, tool calls, entity resolution, code diffs, scheduling, multi-hop ranking, multilingual text, exact posteriors, many-option linking, broad real-language judgements and long documents. Labels come from people or real systems where they exist and are computed by code for the generated rows. The Decision Index 0.2 benchmark is excluded (its datasets by name, and every state checked against its items).
Citation
@software{lod_2026,
title = {lod-stor-4B: a prefill-only option-scoring model for calibrated decisions},
year = {2026},
note = {Qwen/Qwen3-4B-Base + LoRA, independent questions and options, a comparison
layer over the shortlist, proper-scoring fine-tuning, separate confidence
head}
}
- Downloads last month
- 8
Model tree for mrn-dk/lod-stor-4B
Base model
Qwen/Qwen3-4B-BaseEvaluation results
- Accuracy (micro) on Lod held-out test tasksself-reported0.789
- NLL vs soft targets (micro) on Lod held-out test tasksself-reported0.665
- Expected calibration error on Lod held-out test tasksself-reported0.036
- AURC, confidence head on Lod held-out test tasksself-reported0.045