lenabarretta/sharada-multilingual-base

A small encoder that makes typed decisions about text in one forward pass. The options come with the request and what comes back is a probability for each of them: nothing is generated and nothing is parsed. Because the labels are part of the input rather than part of the weights, a label set it has never been trained on still gets an answer.

Code, examples and the training run: https://github.com/LenaBarretta/sharada. The design and the experiments behind it: https://lenatriestounderstand.com/notes/llm/024-rlcr/.

pip install sharada
from sharada import DecisionModel

model = DecisionModel.from_pretrained("lenabarretta/sharada-multilingual-base")

d = model.decide(
    text="My card still hasn't arrived and I ordered it two weeks ago.",
    question="Which team should handle this?",
    options=["billing", "card delivery", "technical support", "account closure"],
)
d.answer, d.confidence, d.probabilities

Fine-tuning on a few hundred of your own labelled examples is the intended use:

report = model.fit(examples)      # holds 20% out, early-stops, then fits a temperature per task
model.save("my-router")

What it is

encoder jhu-clsp/mmBERT-base
parameters 308M
text up to 256 tokens, plus 48 for the question and 12 per option
kinds of question choice (unordered labels), scale (ordered steps), binary (yes or no)
output one probability per option, from a single forward pass
weights float16 on disk, float32 in memory — a storage format, not quantisation
one decision 23.1 ms per question

Three properties hold by construction rather than by training: the order of the options cannot change the answer (every option branch starts at the same position id), an option's score does not depend on which other options are offered (an option reads the text, the question and itself), and the text is read once however many questions are asked of it. The tests in the repository check all three on an untrained model.

How it was trained

663,357 examples from 35 public label sets — intents, topics, review scores, emotion, toxicity, spam and entailment — for 15,000 steps of 32, AdamW at 3e-05 with a cosine schedule, cross entropy loss. The published weights are the ones that measured best on held-out data, at step 15,000 of 15,000 — past that the model stops answering better and only grows more certain. 7 further label sets were held out of training entirely and only measured: arxiv-category, claim-veracity, massive-scenario, medical-pair, poem-tone, subjective, topic-unseen-languages.

In training the options were shuffled, long label sets were often shown as a sampled handful, and each label set was asked through several wordings of its question — so the model reads the options and the question rather than their positions.

What it scores

Trained on is what that label set actually contributed to this run — not its cap, since a small dataset runs out first and a pooled multilingual one is split between its languages; a dash means the model never saw it. Measured on is how many held-out examples the accuracy beside it rests on, and it is worth reading first: a row measured on ninety-six examples moves by a full point when one answer changes.

Measured on held-out examples, with every label offered at once — all 151 intents of clinc, all 77 of banking — because that is what a request actually looks like. One temperature per label set was fitted on the same held-out examples. ECE is the expected calibration error over 15 equal bands: how far the stated probability is from how often it turns out right. Rows marked unseen are label sets kept out of training entirely, never trained on, only measured.

label set answer options kind trained on measured on accuracy log loss ECE T
clinc-intent 151 choice 24,000 480 0.883 0.417 0.035 0.891
banking-intent 77 choice 19,986 480 0.854 0.494 0.034 1.0
massive-intent 59 choice 23,028 480 0.875 0.447 0.047 1.26
question-type-fine 50 choice 10,904 240 0.900 0.287 0.059 1.122
massive-intent-multi 35 choice 72,000 1,440 0.849 0.535 0.032 1.414
fine-emotion 28 choice 18,000 360 0.581 1.319 0.051 1.122
newsgroup 20 choice 14,592 292 0.685 0.930 0.078 1.189
entity-type 14 choice 18,000 360 0.997 0.004 0.003 0.375
forum-topic 10 choice 18,000 360 0.769 0.706 0.042 0.891
topic-multi 7 choice 30,398 660 0.815 0.596 0.044 2.119
question-type 6 choice 10,904 240 0.975 0.084 0.016 1.26
emotion 6 choice 15,000 300 0.913 0.256 0.025 1.414
app-stars 5 scale 15,000 300 0.677 0.867 0.047 1.059
review-stars 5 scale 24,000 480 0.650 0.752 0.045 0.944
sentence-tone 5 scale 15,000 300 0.570 0.998 0.083 1.26
news-section 4 choice 18,000 360 0.933 0.169 0.017 0.944
tweet-emotion 4 choice 6,514 240 0.817 0.432 0.057 1.059
entailment-short 3 scale 18,000 360 0.875 0.342 0.045 0.944
entailment 3 scale 24,000 480 0.815 0.491 0.027 1.189
entailment-multi 3 scale 71,968 1,428 0.764 0.562 0.027 1.26
tweet-sentiment 3 scale 15,000 300 0.680 0.699 0.053 1.26
spam 2 binary 10,034 240 0.992 0.024 0.005 1.059
product-tone 2 scale 15,000 300 0.947 0.126 0.025 0.891
short-verdict 2 scale 12,000 240 0.938 0.183 0.025 0.891
movie-verdict 2 scale 15,000 300 0.933 0.167 0.051 1.059
toxic-comment 2 binary 15,000 300 0.920 0.199 0.039 1.059
answers-question 2 binary 18,000 360 0.892 0.282 0.011 1.0
paraphrase 2 binary 15,000 300 0.883 0.265 0.043 1.414
offensive 2 binary 15,000 300 0.863 0.322 0.032 0.841
same-question 2 binary 18,000 360 0.825 0.380 0.034 1.0
hateful 2 binary 14,989 300 0.810 0.387 0.052 1.122
same-meaning 2 binary 7,336 240 0.787 0.461 0.033 1.498
follows 2 binary 4,980 240 0.779 0.481 0.053 1.414
grammatical 2 binary 15,000 300 0.740 0.504 0.082 1.414
irony 2 binary 5,724 240 0.725 0.567 0.077 0.749
massive-scenario (unseen) 18 choice — 240 0.767 0.692 0.088 1.26
arxiv-category (unseen) 11 choice — 180 0.522 1.441 0.080 1.122
topic-unseen-languages (unseen) 7 choice — 240 0.679 0.912 0.068 1.782
poem-tone (unseen) 4 scale — 96 0.510 1.126 0.205 1.414
claim-veracity (unseen) 4 choice — 180 0.383 1.338 0.092 2.828
medical-pair (unseen) 2 binary — 180 0.750 0.558 0.061 2.0
subjective (unseen) 2 choice — 180 0.461 0.709 0.077 8.0

Overall: accuracy 0.802, log loss 0.518, Brier 0.269, calibration error 0.009 (95% interval [0.008, 0.016]).

The family

checkpoint encoder download
lenabarretta/sharada-base ModernBERT-base, 150M 300 MB the default, and the one to fine-tune
lenabarretta/sharada-large ModernBERT-large, 397M 794 MB a few points better, about a third slower
lenabarretta/sharada-multilingual-base — you are here mmBERT-base, 308M 616 MB the same model over many more languages
lenabarretta/sharada-multilingual-small mmBERT-small, 141M 282 MB narrower body, for throughput rather than for one fast answer

Everything that makes a model what it is lives in the checkpoint: which encoder it wants, how long a text it reads, the temperature fitted for each task. The library reads all of that out of config.json, so a bigger model or a multilingual one is another repository rather than another version of sharada — and from_pretrained loads any of them with the same line.

That is also why a multilingual model cannot be a flag on this one. This encoder is English down to its tokenizer; reading another language means other weights, not another setting.

The probabilities expire

A temperature is fitted on a distribution, not on a model, so it goes stale when the traffic moves. The run's passport.json records what it was fitted on and what should make you fit it again:

from sharada import check_passport
import json

check_passport(json.load(open("passport.json")), recent_examples, model)
# [] when it finds nothing wrong

Fit it again on your own labelled examples before trusting the numbers on your own traffic — model.calibrate(examples) does it in one call, and model.fit(examples) does it for you.

What it will not do

It does not generate: options or nothing. It reads 256 tokens of text, so longer documents need chunking. It has seen the forty-odd languages of its mix and has not been measured on any outside them. And there is no medicine, law or code in its training mix: the two medical label sets are measured only, and verifying public-health claims lands at the majority-class baseline, which is to say it does not work. Those domains need fine-tuning on your own labelled examples. Finally it is small — where an answer needs a fact that is not in the text in front of it, a frontier model wins.

Apache 2.0.

Downloads last month
65
Safetensors
Model size
0.3B params
Tensor type
F16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for lenabarretta/sharada-multilingual-base

Finetuned
(172)
this model