longburn

A 0.66M-parameter byte-level encoder with an attention read-out head, trained to pick the option that shares a cue substring with the context. Published as the reference checkpoint the longburn benchmark was measured against.

Read What this number is worth before quoting the accuracy. It is real, and it is also not measuring the rule the task was built to measure.

The task

Each record pairs a context with 2-6 candidate options and asks for one of three decision primitives:

Primitive Question Label
choice which option an option index
score how strongly an ordinal value in [0, 1]
noul whether at all a boolean over a two-way list

The correct option is the one carrying a cue substring that also occurs in the context, which forces a context-by-option interaction rather than a prior over option position or length.

Architecture

Encoder 2 layers, hidden 128, 4 heads
Vocabulary byte-level, 261 entries; max context 256 bytes
Attention sliding window with a global layer every 2
Head 1 refinement layer, 4 heads, rank-32 option attention, plus a contiguous-byte-run read over the options
Decision types 3, up to 32 options per record
Parameters 0.66M
Training early stopping at epoch 8 of 40 requested
Best validation loss 0.085706

The weights are a Burn named-tensor record: not safetensors, and not a torch.save pickle. They load only under the exact LongburnModelConfig that produced them, which checkpoint.json carries and config.toml reproduces.

Files

File What it is
model.weights the Burn record. The CLI takes a base path and appends .weights itself
calibrator.json temperature scaler fitted on a held-out-phrasing split. Apply it β€” see Calibration
checkpoint.json training metadata: architecture, optimiser and trainer configs, epoch counters. Authoritative for the architecture
config.toml that architecture as a longburn config file, for --config
PROVENANCE.md the training recipe, and the numbers a retrain has to reproduce
MANIFEST.sha256 sha256 of every file here; check with sha256sum -c MANIFEST.sha256

Usage

from huggingface_hub import hf_hub_download

repo = "Akagi201/longburn"
base = hf_hub_download(repo, "model.weights").removesuffix(".weights")
config = hf_hub_download(repo, "config.toml")

Score a single request:

longburn infer --config config.toml --model "$base" \
    --calibrator calibrator.json \
    --context "the deployment failed, roll it back" \
    --options "roll back,scale up,ignore" \
    --decision-type choice

Score a labelled JSONL file, the way the benchmark does:

longburn eval --config config.toml --model "$base" \
    --calibrator calibrator.json --data test.jsonl

What this number is worth

Measurement Value
Train top-1 99.96%
Held-out test top-1 81.14% (82.40% choice, 75.40% noul)
Best constant-index baseline 33.10%
Expected calibration error, test split 0.0611

Top-1 is scored on the 562 of 750 held-out records whose decision type has a discrete answer; score records have an ordinal label and contribute to the proper scoring loss only.

Calibration

This checkpoint ships a fitted temperature scaler, and it is not optional: raw, the model reports 87.26% confidence on the held-out split while being right 80.07%.

calibrator.json carries one temperature per (decision type, option count), fitted on one half of the held-out-surface-form split. The other half is what training early-stops on. The two share no rows, so the temperature is not fitted on data that chose the checkpoint β€” but they do share vocabulary by construction, since one held-out vocabulary serves both jobs.

Held-out split Mean confidence Top-1 ECE (15 bins)
raw 87.26% 80.07% 0.0846
calibrated 83.24% 80.07% 0.0611

Fitted temperatures: 0.5-5. The choice groups fit below 1, which is the number worth noticing: off-distribution this model is mildly underconfident rather than badly overconfident, where an earlier release needed temperatures of 12 and beyond to say anything sensible about its confidence at all. A discriminative model does not need its logits squashed to report how often it is right. The noul group still fits above 1; that is the abstention decision, which this benchmark does not score. On the training distribution the model is nearly calibrated, so any remaining shift is a generalisation effect rather than a broken objective β€” the training loss is already a strictly proper scoring rule.

This is a mitigation, not a fix, and the distinction is the point. Three things the table does not say on its own:

  • Accuracy is identical in both rows and cannot differ: temperature scaling is monotone within a row, so it moves probability between options and never reorders them.
  • The residual 0.0611 is a discrimination failure rather than a scale error. At 83.24% confidence against 80.07% correct the model is close to calibrated on average, but ECE averages: the rows it gets wrong on held-out phrasings it gets wrong confidently, and no temperature expresses that, because flattening further destroys the ranking along with the confidence. It is the same generalisation gap, seen from the confidence side.
  • The same calibrator makes the training split worse (ECE 0.0064 β†’ 0.0610), because a temperature fitted for shifted data over-flattens data the model is good at. One temperature is being asked to serve two regimes.

Apply it for off-distribution serving. Do not read it as a calibrated model.

Two caveats that outrank the ranking

Both are about the model rather than the benchmark.

  1. The model performs the rule on most rows, not all of them. Decorrelating the contexts from their option lists costs 47.5 points and what survives lands on the 33.10% constant-index baseline, so most of its score is context-driven. On the choice rows β€” the only decision type where the rule is substring matching β€” it scores 82.40% against a 33.10% prior, and with the contexts scrambled it falls below that prior. It is reading the context and matching bytes against it; it is not right every time, and it does not transfer perfectly to phrasings it has never seen.

    The benchmark itself is sound, and just bench-data refuses to write a split where that is not true: the rule oracle scores 100% on every split, every test pair is unseen in train, and an options-only probe that never reads a context scores below the constant-index prior.

    Two leaks were closed to get here, which is why the accuracy above is lower than earlier releases of this card. First the distractors came from a vocabulary disjoint from the answers, so the option list answered the question on its own at 100%. Then, with that closed, early stopping ran against a random row split β€” which contains the same context templates as training β€” so the selected epoch rewarded memorising "template family β†’ answer form". This checkpoint early-stops on a split holding out context templates and channels instead.

    A third defect was in the architecture rather than the data, and it is the one worth knowing about if you read the source. The head scored each option by a max over individual bytes β€” a disjunction, "some byte matches somewhere" β€” which any sentence long enough to repeat a letter satisfies by accident. It now scores contiguous runs of bytes, divided by sqrt(length), which is what "these bytes occur here, in order" means. That change alone moved choice from 41.6% to 82.40%.

  2. Apply calibrator.json before trusting the confidence off-distribution. It takes the expected calibration error from 0.0846 to 0.0611 and fixes nothing else β€” not one decision changes, because temperature scaling is monotone within a row. See Calibration for why the residual is a discrimination failure rather than a scale error.

The generalisation gap

99.96% on train against 81.14% on held-out cue phrasings, templates and answer forms. The model fits its data and does not generalise across surface form, and on the training distribution it is nearly calibrated β€” so the gap is a generalisation failure, not a calibration one.

Train accuracy is the figure that moved when the training protocol was fixed: it fell, while held-out accuracy rose. That is the trade being made deliberately. A model that fits its training templates to 91% and scores 38% on a template it has not seen has not learned the rule, and reporting the 91% as the model's quality would be the more flattering and less useful number.

Reproducing this checkpoint

PROVENANCE.md, shipped alongside these weights, records the exact training command.

The measurements above were recorded against longburn commit fc0bedb; this artifact was staged from fc0bedb-dirty. A retrain that lands far from these numbers means something changed β€” check the dataset SHA-256 in the benchmark write-up first, since the generator is deterministic.

License

Apache-2.0. See the license section of the longburn repository.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support