Configuration Parsing Warning:In adapter_config.json: "peft.base_model_name_or_path" must be a string

D1A-E4B v0.4

D1A is a small open decision model: one document and a set of typed questions in (yes/no, choice, score), a calibrated probability for every option out, in one forward pass, with no generated text. It speaks the TypeSafe System One API (POST /v1/systemone), so the TypeSafe SDK and existing clients work against it.

v0.4 is a second pull-request labeling round on top of v0.3: better severity and overall labels in English and Japanese, a blast-radius answer that no longer leans to "broad", and confident (p ≥ 0.7) on more labels. It costs 1–3 points on the earlier general, Japanese and routing skills; see Results. v0.3 stays available for those uses.

What's inside

A LoRA adapter (rank 16 on the attention and MLP projections: q, k, v, o, gate, up, down) and a small pointer head (256-dimensional) on google/gemma-4-E4B at revision 411aa17b. Each version continues training the previous one, so v0.4 contains every stage below:

Version Stage Training data Settings
v0.1 General decisions decision-v7 suite: ~12,500 records from public classification datasets plus generated policy and rule data 2 epochs, lr 5e-5
(v0.2) Kev's later skills dates and missing evidence (1,425), hard decisions, developer-tool decisions and long documents (hard-v1, devtools-v1, documents-v1 training partitions, states up to 4,096 tokens), decision-v7 replay 4,000 1 epoch, lr 2e-5
v0.2 Japanese and agent routing Japanese decisions from JGLUE (JNLI, JCommonsenseQA, JSTS; 6,000), agent-factory and model-routing questions (2,745), decision-v7 replay 3,000 1 epoch, lr 5e-5
v0.3 Pull-request labeling 3,500 English and 768 Japanese pull requests labeled with change type, blast radius and severity (closed PRs of NousResearch/hermes-agent, MIT; Japanese copies machine-translated; severity follows the labeling job's rules for docs-only and dependency PRs), replay: decision-v7 1,500, JGLUE 500, routing 300 1 epoch, lr 5e-5, documents up to 1,536 tokens
v0.4 Pull-request labeling, round 2 the 4,063 English PRs v0.3 did not see, its 2,396 rare-label PRs (blast radius, P0, P1, P4) again, 2,185 more PRs with a blast-radius label (blast and severity questions only; all older than every evaluation PR), the 768 Japanese PRs again; replay: decision-v7 1,500, JGLUE 500, routing 300 (11,540 records after dropping 172 too long) from v0.3, 1 epoch, lr 5e-5, documents up to 1,536 tokens, 1,443 steps on one L4

Calibration: a single temperature, T = 1.59, fitted on pooled held-out rows (decision-v7 calibration split and PR-labeling development set; calibration error 0.066 before, 0.019 after, out of fold), stored in head.pt.

Results

Every row scored through the same HTTP client (scripts/compare_systemone.py) on the MLX 8-bit builds, held-out data only:

Set v0.3 v0.4
PR labels, English (953 PRs): change type / severity / blast radius 89% / 70% / 54% 88% / 78% / 51%
PR labels, English, all labels 78% 81%
PR labels, English, labels where p ≥ 0.7 93% right on 57% of labels 92% right on 65% of labels
PR labels, Japanese (92 PRs) 70% 74%
17 real open PRs from the owner's other repositories (labels drafted by Claude) 61% 69%
JGLUE development (500, Japanese) 85% 82%
Held-out transfer-v4 development (764, sources never trained on) 71% 69%
Agent factory development (640 questions) 92% 91%
Model routing: generic development (270) / hand-labelled (45) 99% / 100% 97% / 100%

Blast radius: v0.3 answered "broad" for 81 of 139 test PRs; v0.4 spreads its answers (33 broad, 67 moderate, 39 contained) and finds more contained (20 of 42, was 14) and moderate (28 of 44, was 20) PRs, fewer broad ones (23 of 44, was 41). Its accuracy is about the same; it is less biased.

The questions it was trained for

PR labeling works best with these questions, worded exactly so, over a document of the form title / author / stats / body / files (one - status path +added/-deleted line per file):

  • type (choice): "Primary change type from files and body, not the title prefix." Options: type/bug, type/docs, type/feature, type/perf, type/refactor, type/security, type/test.
  • blast (choice): "How far a mistake in this PR spreads in production." Options: review:blast-contained (one module), -moderate (one subsystem), -broad (shared helper or config), -massive (auth, permissions, or all paths).
  • sev (choice): "How serious the problem this PR addresses is — not the risk of merging the diff as-is." Options P0 to P4.

The PR-labeler recipe has the exact wording, the document builder and the training and scoring scripts. Other questions about other documents work as in v0.2.

Run it

The checkpoint runs on any device; no retraining per device:

pip install "d1a[serve] @ git+https://github.com/jonpol01/d1a"
python -m d1a.serve --run JohnP1/d1a-e4b@v0.4 --port 8009        # NVIDIA GPU, CPU, or Apple GPU (PyTorch)
python -m d1a.serve --run JohnP1/d1a-e4b-mlx-q8@v0.4 --port 8009  # Apple Silicon: the 8-bit MLX build (~6 GB)

or in-process: from d1a import D1A; D1A.load("JohnP1/d1a-e4b@v0.4").decide(state, questions).

Limits

  • Blast radius is still the weakest label (51% on the English test) and has never picked "massive" (0 of 9; only 87 such PRs exist in the source project). A "needs a human" rule should not rest on it alone.
  • Severity is still weakest at the urgent end: on the English test it finds 13 of 38 P0 and 17 of 47 P1 PRs (v0.3: 16 and 15), against 505 of 564 P3 and 75 of 92 P4.
  • The earlier skills are 1–3 points below v0.3 (table above). For general decisions, Japanese JGLUE-style questions or routing alone, v0.3 is the better choice.
  • The PR data comes from one large open-source project; label conventions follow its maintainers except where the labeling job's rules override them.
  • Questions are in English; the documents can be English or Japanese.

Versions

Tag What it adds
v0.1 general decisions (decision-v7, 2 epochs, calibrated)
v0.2 Kev's later skills, Japanese, agent routing
v0.3 pull-request labeling, English and Japanese
v0.4 pull-request labeling round 2: severity, Japanese, less biased blast radius

Older tag names (v0.1-2epoch, v0.1.1-2epoch-calibrated, v0.2-hybrid) still work.

License and data

Apache-2.0. Base model: Gemma 4 by Google (Apache-2.0). Code: github.com/jonpol01/d1a, built on Kev by Jared Palmer (Apache-2.0). Japanese decision data derived from JGLUE by Yahoo Japan Corporation and Waseda University (CC BY-SA 4.0). Pull-request data from NousResearch/hermes-agent (MIT); the labeled PR dataset itself is private.

Downloads last month
39
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for JohnP1/d1a-e4b

Adapter
(20)
this model
Quantizations
1 model