udaykamath/OpenSystemOne-ModernBERT-base-v0.1

An open System One classifier: typed questions about a piece of context, answered with calibrated probabilities. Yes/no (Noul), pick-one (Choice) and ordered-scale (Score) questions, many per call.

Inspired by the public description of TypeSafe AI's Jev. Not affiliated with TypeSafe AI and not a copy of Jev.

s1 = SystemOne.from_pretrained("udaykamath/OpenSystemOne-ModernBERT-base-v0.1")
s1({"message": "I was charged twice and need a refund today."},
   {"urgent": Noul("`message` communicates time pressure."),
    "intent": Choice("What does the customer want?", {"refund": "money back", "technical_help": "a bug fixed"})})
Design cross-encoder: state and question read together
Backbone tasksource/ModernBERT-base-nli
Context 512 tokens of state, 128 of question
Temperature 1.402 (fitted on held-aside training-mix tasks)

Held-out results (zero-shot)

task acc macro_F1 ECE
IMDb (noul) 0.925 0.925 0.028
IMDb flipped (noul) 0.921 0.921 0.019
IMDb (choice) 0.940 0.940 0.025
AG News (choice) 0.870 0.869 0.076
TweetEval-hate (noul) 0.503 0.457 0.250
SST-5 (score) 0.487 0.492 0.159
SMS spam, 4 wordings (noul) 0.818 0.763 0.030
Banking77, 8 options (choice) 0.844 0.843 0.162
Yahoo Answers topics (choice) 0.674 0.667 0.114

Robustness

  • negation: mean |P(liked)+P(disliked)-1|: 0.088
  • negation: correlation: -0.929
  • paraphrase: mean std of P(yes): 0.030
  • paraphrase: worst-wording accuracy: 0.880
  • paraphrase: best-wording accuracy: 0.920
  • option order: same answer under 4 rotations: 0.963
  • none trap: picks 'none' when the answer is missing: 0.853
  • none decoy: accuracy with 'none' added: 0.777
  • labels: accuracy with bare names: 0.873
  • labels: accuracy with descriptions: 0.860
  • context: IMDb accuracy at 128 tokens: 0.860
  • context: IMDb accuracy at 256 tokens: 0.897
  • context: IMDb accuracy at 512 tokens: 0.913

Training data and licences

see repository

Some training sources carry non-commercial terms. See DATA_LICENSES.md in the code repository before commercial use.

Limitations

  • Each question is a separate pass over the state, so 16 questions cost about 16 times one.
  • Offensive and hateful text are not well separated: on TweetEval-hate the model ranks hate reasonably (AUC about 0.7) but says yes far more often than the labels do.
  • Literal cues (explicit urgency, dates, numbers) can be missed.
  • Calibration was fitted on the training mix and can drift on a new domain. Re-fit the temperature on your own sample.
  • English only.
Downloads last month
11
Safetensors
Model size
0.2B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for udaykamath/OpenSystemOne-ModernBERT-base-v0.1

Finetuned
(9)
this model