gev-e2b — Gemma 4 E2B decision model

State and typed questions in; one answer and a probability distribution per question out. Gev scores the supplied options rather than generating text. It combines a LoRA adapter on google/gemma-4-E2B with a separately trained pointer head.

The pointer head (pointer.safetensors) is required: loading the PEFT adapter alone does not produce Gev decisions. The Gev package loads the base model, adapter, pointer head, tokenizer markers, and serving temperature. Its source repository is currently private; access to it is needed to run inference.

  • Trained-source test: 82.9% accuracy on 1,200 clean questions.
  • New-source test: 62.5% accuracy on 656 clean questions.

Use

With access to the Gev source, run uv sync --locked in its checkout (add --extra mlx on Apple Silicon). Save this as request.json:

{"state":"Order #1 arrived damaged.","questions":{"route":{"type":"choice","instructions":"Which team handles this?","criteria":{"billing":"Payments","support":"Product issues"}}}}
uv run gev predict onatm/gev-e2b --input request.json

The response includes questions.route.answer, questions.route.probabilities (one per option), and the serving temperature. Choice, yes/no (noul), and ordinal (score) questions are supported. A Hugging Face text-generation or PEFT-only pipeline cannot serve this model.

Evaluation

Clean-question scores from the saved reports. Accuracy is at raw T=1 (temperature scaling does not change the winning answer); Brier and ECE are lower-is-better. A dash means calibrated metrics were not recorded for that split.

The served columns use the saved temperature T=1.6245.

Suite / split Clean n Accuracy Brier raw Brier served ECE raw ECE served
decision-v7/development 1,264 0.7975 0.2847 - 0.0628 -
transfer-v4/development 656 0.6113 0.4990 0.4735 0.1163 0.0522
decision-v7/test 1,200 0.8292 0.2457 0.2336 0.0661 0.0170
transfer-v4/test 656 0.6250 0.4680 0.4491 0.1136 0.0571

decision-v7 contains held-out questions from the training source families; transfer-v4 contains new sources and held-out policy structures. Development was for model selection; these reports describe seed 0. Temperature was fitted on the separate decision-v7/calibration split, not on either test split.

The decision-v7 test report predates the temperature fit. Its served metrics were computed afterward from saved raw logits, without rerunning inference or fitting on test.

Known limits

This is one seed, not a multi-seed study. The new-source test is substantially harder than the trained-source test; evaluate on your own decisions before use.

  • For changed-answer contrastive pairs, both answers were correct in 12/64 pairs.
  • With a none-of-the-above option present, 14/36 questions were correct.

Model and provenance

  • Base: google/gemma-4-E2B at revision d29ff6b45f081a49ee2733a859c9c9c2d95d1a6f; frozen base weights are not in this repo.
  • Architecture: LoRA rank 16, alpha 32 on the text decoder and a 256-wide pointer head. Each question is scored independently.
  • Recipe: seed 0, 2 epochs, 12,576 training records, 3,144 steps, MLX/BF16. The checkpoint stores the fitted temperature in gev.json.
  • Training data SHA-256: 7ed5254b5cb5291baefaceb09edf7e13110258211518c8038f4a12c11bd628ad; each saved evaluation report also records its split's SHA-256.
  • Training suite: jaredpalmer/kev-suites (decision-v7, pinned and verified by hash).

Training uses option permutation, none-of-the-above and distractor augmentation, and contrastive pairs. The architecture follows Jev's Architecture Unmasked; the data and evaluation protocol are adapted from Kev. The detailed run reports and code are in the private Gev repository.

License

The adapter and pointer head are apache-2.0; check the separately loaded base model and dataset licenses as well.

Downloads last month
20
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for onatm/gev-e2b

Adapter
(37)
this model

Dataset used to train onatm/gev-e2b

Evaluation results

  • accuracy on decision-v7/test (clean questions)
    self-reported
    0.829
  • Brier (as served) on decision-v7/test (clean questions)
    self-reported
    0.234
  • ECE (as served) on decision-v7/test (clean questions)
    self-reported
    0.017
  • accuracy on transfer-v4/test (clean questions)
    self-reported
    0.625
  • Brier (as served) on transfer-v4/test (clean questions)
    self-reported
    0.449
  • ECE (as served) on transfer-v4/test (clean questions)
    self-reported
    0.057