lenabarretta/sharada-multilingual-base
A small encoder that makes typed decisions about text in one forward pass. The options come with the request and what comes back is a probability for each of them: nothing is generated and nothing is parsed. Because the labels are part of the input rather than part of the weights, a label set it has never been trained on still gets an answer.
Code, examples and the training run: https://github.com/LenaBarretta/sharada. The design and the experiments behind it: https://lenatriestounderstand.com/notes/llm/024-rlcr/.
pip install sharada
from sharada import DecisionModel
model = DecisionModel.from_pretrained("lenabarretta/sharada-multilingual-base")
d = model.decide(
text="My card still hasn't arrived and I ordered it two weeks ago.",
question="Which team should handle this?",
options=["billing", "card delivery", "technical support", "account closure"],
)
d.answer, d.confidence, d.probabilities
Fine-tuning on a few hundred of your own labelled examples is the intended use:
report = model.fit(examples) # holds 20% out, early-stops, then fits a temperature per task
model.save("my-router")
What it is
| encoder | jhu-clsp/mmBERT-base |
| parameters | 308M |
| text | up to 256 tokens, plus 48 for the question and 12 per option |
| kinds of question | choice (unordered labels), scale (ordered steps), binary (yes or no) |
| output | one probability per option, from a single forward pass |
| weights | float16 on disk, float32 in memory — a storage format, not quantisation |
| one decision | 23.1 ms per question |
Three properties hold by construction rather than by training: the order of the options cannot change the answer (every option branch starts at the same position id), an option's score does not depend on which other options are offered (an option reads the text, the question and itself), and the text is read once however many questions are asked of it. The tests in the repository check all three on an untrained model.
How it was trained
663,357 examples from 35 public label sets — intents, topics, review scores, emotion, toxicity, spam and entailment — for 15,000 steps of 32, AdamW at 3e-05 with a cosine schedule, cross entropy loss. The published weights are the ones that measured best on held-out data, at step 15,000 of 15,000 — past that the model stops answering better and only grows more certain. 7 further label sets were held out of training entirely and only measured: arxiv-category, claim-veracity, massive-scenario, medical-pair, poem-tone, subjective, topic-unseen-languages.
In training the options were shuffled, long label sets were often shown as a sampled handful, and each label set was asked through several wordings of its question — so the model reads the options and the question rather than their positions.
What it scores
Trained on is what that label set actually contributed to this run — not its cap, since a small dataset runs out first and a pooled multilingual one is split between its languages; a dash means the model never saw it. Measured on is how many held-out examples the accuracy beside it rests on, and it is worth reading first: a row measured on ninety-six examples moves by a full point when one answer changes.
Measured on held-out examples, with every label offered at once — all 151 intents of clinc, all 77 of banking — because that is what a request actually looks like. One temperature per label set was fitted on the same held-out examples. ECE is the expected calibration error over 15 equal bands: how far the stated probability is from how often it turns out right. Rows marked unseen are label sets kept out of training entirely, never trained on, only measured.
| label set | answer options | kind | trained on | measured on | accuracy | log loss | ECE | T |
|---|---|---|---|---|---|---|---|---|
| clinc-intent | 151 | choice | 24,000 | 480 | 0.883 | 0.417 | 0.035 | 0.891 |
| banking-intent | 77 | choice | 19,986 | 480 | 0.854 | 0.494 | 0.034 | 1.0 |
| massive-intent | 59 | choice | 23,028 | 480 | 0.875 | 0.447 | 0.047 | 1.26 |
| question-type-fine | 50 | choice | 10,904 | 240 | 0.900 | 0.287 | 0.059 | 1.122 |
| massive-intent-multi | 35 | choice | 72,000 | 1,440 | 0.849 | 0.535 | 0.032 | 1.414 |
| fine-emotion | 28 | choice | 18,000 | 360 | 0.581 | 1.319 | 0.051 | 1.122 |
| newsgroup | 20 | choice | 14,592 | 292 | 0.685 | 0.930 | 0.078 | 1.189 |
| entity-type | 14 | choice | 18,000 | 360 | 0.997 | 0.004 | 0.003 | 0.375 |
| forum-topic | 10 | choice | 18,000 | 360 | 0.769 | 0.706 | 0.042 | 0.891 |
| topic-multi | 7 | choice | 30,398 | 660 | 0.815 | 0.596 | 0.044 | 2.119 |
| question-type | 6 | choice | 10,904 | 240 | 0.975 | 0.084 | 0.016 | 1.26 |
| emotion | 6 | choice | 15,000 | 300 | 0.913 | 0.256 | 0.025 | 1.414 |
| app-stars | 5 | scale | 15,000 | 300 | 0.677 | 0.867 | 0.047 | 1.059 |
| review-stars | 5 | scale | 24,000 | 480 | 0.650 | 0.752 | 0.045 | 0.944 |
| sentence-tone | 5 | scale | 15,000 | 300 | 0.570 | 0.998 | 0.083 | 1.26 |
| news-section | 4 | choice | 18,000 | 360 | 0.933 | 0.169 | 0.017 | 0.944 |
| tweet-emotion | 4 | choice | 6,514 | 240 | 0.817 | 0.432 | 0.057 | 1.059 |
| entailment-short | 3 | scale | 18,000 | 360 | 0.875 | 0.342 | 0.045 | 0.944 |
| entailment | 3 | scale | 24,000 | 480 | 0.815 | 0.491 | 0.027 | 1.189 |
| entailment-multi | 3 | scale | 71,968 | 1,428 | 0.764 | 0.562 | 0.027 | 1.26 |
| tweet-sentiment | 3 | scale | 15,000 | 300 | 0.680 | 0.699 | 0.053 | 1.26 |
| spam | 2 | binary | 10,034 | 240 | 0.992 | 0.024 | 0.005 | 1.059 |
| product-tone | 2 | scale | 15,000 | 300 | 0.947 | 0.126 | 0.025 | 0.891 |
| short-verdict | 2 | scale | 12,000 | 240 | 0.938 | 0.183 | 0.025 | 0.891 |
| movie-verdict | 2 | scale | 15,000 | 300 | 0.933 | 0.167 | 0.051 | 1.059 |
| toxic-comment | 2 | binary | 15,000 | 300 | 0.920 | 0.199 | 0.039 | 1.059 |
| answers-question | 2 | binary | 18,000 | 360 | 0.892 | 0.282 | 0.011 | 1.0 |
| paraphrase | 2 | binary | 15,000 | 300 | 0.883 | 0.265 | 0.043 | 1.414 |
| offensive | 2 | binary | 15,000 | 300 | 0.863 | 0.322 | 0.032 | 0.841 |
| same-question | 2 | binary | 18,000 | 360 | 0.825 | 0.380 | 0.034 | 1.0 |
| hateful | 2 | binary | 14,989 | 300 | 0.810 | 0.387 | 0.052 | 1.122 |
| same-meaning | 2 | binary | 7,336 | 240 | 0.787 | 0.461 | 0.033 | 1.498 |
| follows | 2 | binary | 4,980 | 240 | 0.779 | 0.481 | 0.053 | 1.414 |
| grammatical | 2 | binary | 15,000 | 300 | 0.740 | 0.504 | 0.082 | 1.414 |
| irony | 2 | binary | 5,724 | 240 | 0.725 | 0.567 | 0.077 | 0.749 |
| massive-scenario (unseen) | 18 | choice | — | 240 | 0.767 | 0.692 | 0.088 | 1.26 |
| arxiv-category (unseen) | 11 | choice | — | 180 | 0.522 | 1.441 | 0.080 | 1.122 |
| topic-unseen-languages (unseen) | 7 | choice | — | 240 | 0.679 | 0.912 | 0.068 | 1.782 |
| poem-tone (unseen) | 4 | scale | — | 96 | 0.510 | 1.126 | 0.205 | 1.414 |
| claim-veracity (unseen) | 4 | choice | — | 180 | 0.383 | 1.338 | 0.092 | 2.828 |
| medical-pair (unseen) | 2 | binary | — | 180 | 0.750 | 0.558 | 0.061 | 2.0 |
| subjective (unseen) | 2 | choice | — | 180 | 0.461 | 0.709 | 0.077 | 8.0 |
Overall: accuracy 0.802, log loss 0.518, Brier 0.269, calibration error 0.009 (95% interval [0.008, 0.016]).
The family
| checkpoint | encoder | download | |
|---|---|---|---|
lenabarretta/sharada-base |
ModernBERT-base, 150M | 300 MB | the default, and the one to fine-tune |
lenabarretta/sharada-large |
ModernBERT-large, 397M | 794 MB | a few points better, about a third slower |
lenabarretta/sharada-multilingual-base — you are here |
mmBERT-base, 308M | 616 MB | the same model over many more languages |
lenabarretta/sharada-multilingual-small |
mmBERT-small, 141M | 282 MB | narrower body, for throughput rather than for one fast answer |
Everything that makes a model what it is lives in the checkpoint: which encoder it wants, how long a
text it reads, the temperature fitted for each task. The library reads all of that out of
config.json, so a bigger model or a multilingual one is another repository rather than another
version of sharada — and from_pretrained loads any of them with the same line.
That is also why a multilingual model cannot be a flag on this one. This encoder is English down to its tokenizer; reading another language means other weights, not another setting.
The probabilities expire
A temperature is fitted on a distribution, not on a model, so it goes stale when the traffic moves. The
run's passport.json records what it was fitted on and what should make you fit it again:
from sharada import check_passport
import json
check_passport(json.load(open("passport.json")), recent_examples, model)
# [] when it finds nothing wrong
Fit it again on your own labelled examples before trusting the numbers on your own traffic —
model.calibrate(examples) does it in one call, and model.fit(examples) does it for you.
What it will not do
It does not generate: options or nothing. It reads 256 tokens of text, so longer documents need chunking. It has seen the forty-odd languages of its mix and has not been measured on any outside them. And there is no medicine, law or code in its training mix: the two medical label sets are measured only, and verifying public-health claims lands at the majority-class baseline, which is to say it does not work. Those domains need fine-tuning on your own labelled examples. Finally it is small — where an answer needs a fact that is not in the text in front of it, a frontier model wins.
Apache 2.0.
- Downloads last month
- 65
Model tree for lenabarretta/sharada-multilingual-base
Base model
jhu-clsp/mmBERT-base