Add Typed Decisions benchmark results

#2
by codelion - opened
MLX Community org

Adds this model's scores on Typed Decisions as .eval_results/typed-decisions.yaml, so they are listed on the benchmark's Hub leaderboard.

Accuracy KL from gold (lower is better) Brier (lower is better) ECE (lower is better)
0.7185 0.1987 0.106 0.1185

General model (any question schema at request time), not fitted by us on this benchmark's train split. Scored on the full 400-case test split (2,000 decisions): one request per case with all five questions, text only, the benchmark's own metrics, MLX on an Apple M3 Max.

Once merged, the model appears on the leaderboard. Quantizations and other derived models are hidden by default; turn the base-model switch off to see them. While this PR is open the Hub marks the scores as community-provided.

If you would rather not list the model, or a number looks wrong, close this PR or tell us and we will correct it. Scored by codelion (OptiQ).

lucataco changed pull request status to merged

Sign up or log in to comment