udaykamath/OpenSystemOne-ModernBERT-base-v0.1
An open System One classifier: typed questions about a piece of context, answered with calibrated probabilities. Yes/no (Noul), pick-one (Choice) and ordered-scale (Score) questions, many per call.
Inspired by the public description of TypeSafe AI's Jev. Not affiliated with TypeSafe AI and not a copy of Jev.
s1 = SystemOne.from_pretrained("udaykamath/OpenSystemOne-ModernBERT-base-v0.1")
s1({"message": "I was charged twice and need a refund today."},
{"urgent": Noul("`message` communicates time pressure."),
"intent": Choice("What does the customer want?", {"refund": "money back", "technical_help": "a bug fixed"})})
| Design | cross-encoder: state and question read together |
| Backbone | tasksource/ModernBERT-base-nli |
| Context | 512 tokens of state, 128 of question |
| Temperature | 1.402 (fitted on held-aside training-mix tasks) |
Held-out results (zero-shot)
| task | acc | macro_F1 | ECE |
|---|---|---|---|
| IMDb (noul) | 0.925 | 0.925 | 0.028 |
| IMDb flipped (noul) | 0.921 | 0.921 | 0.019 |
| IMDb (choice) | 0.940 | 0.940 | 0.025 |
| AG News (choice) | 0.870 | 0.869 | 0.076 |
| TweetEval-hate (noul) | 0.503 | 0.457 | 0.250 |
| SST-5 (score) | 0.487 | 0.492 | 0.159 |
| SMS spam, 4 wordings (noul) | 0.818 | 0.763 | 0.030 |
| Banking77, 8 options (choice) | 0.844 | 0.843 | 0.162 |
| Yahoo Answers topics (choice) | 0.674 | 0.667 | 0.114 |
Robustness
- negation: mean |P(liked)+P(disliked)-1|: 0.088
- negation: correlation: -0.929
- paraphrase: mean std of P(yes): 0.030
- paraphrase: worst-wording accuracy: 0.880
- paraphrase: best-wording accuracy: 0.920
- option order: same answer under 4 rotations: 0.963
- none trap: picks 'none' when the answer is missing: 0.853
- none decoy: accuracy with 'none' added: 0.777
- labels: accuracy with bare names: 0.873
- labels: accuracy with descriptions: 0.860
- context: IMDb accuracy at 128 tokens: 0.860
- context: IMDb accuracy at 256 tokens: 0.897
- context: IMDb accuracy at 512 tokens: 0.913
Training data and licences
see repository
Some training sources carry non-commercial terms. See DATA_LICENSES.md in the code repository before commercial use.
Limitations
- Each question is a separate pass over the state, so 16 questions cost about 16 times one.
- Offensive and hateful text are not well separated: on TweetEval-hate the model ranks hate reasonably (AUC about 0.7) but says yes far more often than the labels do.
- Literal cues (explicit urgency, dates, numbers) can be missed.
- Calibration was fitted on the training mix and can drift on a new domain. Re-fit the temperature on your own sample.
- English only.
- Downloads last month
- 11
Model tree for udaykamath/OpenSystemOne-ModernBERT-base-v0.1
Base model
answerdotai/ModernBERT-base Finetuned
tasksource/ModernBERT-base-nli