Independent run: our MNLI 0.7667 against your published 0.926

#3
by saidutta69 - opened

Following up on our independent run of decider-4b - a large gap on MNLI we cannot yet explain.

Setup. Revision 533964dae8be, one sealed manifest, seed 42, 1,190 cases / 1,240 scored evaluation decisions across 9 typed-decision suites, bf16 on one T4. Same harness and same pinned revision as our decider-2b post.

What we measured (overall 0.8411):

suite ours your card
MNLI 0.7667 0.926
our moderation suite 0.8790 -
our guardrails suite 0.8333 -

The gap. Both are choice questions over MNLI. Ours is 0.1593 below yours. That is far outside anything our label grouping or option-count could explain, so we do not think we can attribute it to task shape.

What we suspect on our side, most likely first.

  1. Prompt and option verbalisation. We declare the label set per case and score with per_option_conditional_logprob. If your harness instead renders options as an enumerated list and reads a single next-token distribution, the two are not measuring the same function, and MNLI is where that diverges most for us.
  2. Scoring path. Our choice scoring for decider-4b went through the vendor-default readout with no calibration applied.
  3. Split. Our manifest pins case IDs and order indices, so we can diff row-by-row instead of arguing in aggregate.

What would help. The exact MNLI prompt, label set and answer-extraction rule behind your 0.926 figure. We will re-run and publish whichever number holds, including if yours is right and our harness is the bug. We would rather correct our record than leave a 0.17 gap unexplained.

Full data, per-suite matrices and per-row predictions: https://sysone.sdad.pro/

Correction to the setup line above. Two fields in my post were wrong, both mine, neither affecting the MNLI comparison:

field stated above correct
revision 533964dae8be eb5fbdfc9448
scoring path no calibration applied vendor-default readout, dtype bf16
our moderation suite 0.8790 0.9444
our guardrails suite 0.8333 0.8333

The revision is the one to care about: it was transcribed from the wrong row of our results, so the
checkpoint I named was not the checkpoint I ran. eb5fbdfc9448 is the pinned revision from
decider-4b-t4-d0. MNLI remains 0.7667 overall 0.8411, and that is unchanged.

Apologies for the noise. Source of truth for every field is the run directory named above.

Sign up or log in to comment