Instructions to use JohnP1/d1a-e4b with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use JohnP1/d1a-e4b with PEFT:
Task type is invalid.
- Notebooks
- Google Colab
- Kaggle
Configuration Parsing Warning:In adapter_config.json: "peft.base_model_name_or_path" must be a string
D1A-E4B v0.4
D1A is a small open decision model: one document and a set of typed questions in (yes/no, choice, score), a calibrated probability for every option out, in one forward pass, with no generated text. It speaks the TypeSafe System One API (POST /v1/systemone), so the TypeSafe SDK and existing clients work against it.
v0.4 is a second pull-request labeling round on top of v0.3: better severity and overall labels in English and Japanese, a blast-radius answer that no longer leans to "broad", and confident (p ≥ 0.7) on more labels. It costs 1–3 points on the earlier general, Japanese and routing skills; see Results. v0.3 stays available for those uses.
What's inside
A LoRA adapter (rank 16 on the attention and MLP projections: q, k, v, o, gate, up, down) and a small pointer head (256-dimensional) on google/gemma-4-E4B at revision 411aa17b. Each version continues training the previous one, so v0.4 contains every stage below:
| Version | Stage | Training data | Settings |
|---|---|---|---|
| v0.1 | General decisions | decision-v7 suite: ~12,500 records from public classification datasets plus generated policy and rule data | 2 epochs, lr 5e-5 |
| (v0.2) | Kev's later skills | dates and missing evidence (1,425), hard decisions, developer-tool decisions and long documents (hard-v1, devtools-v1, documents-v1 training partitions, states up to 4,096 tokens), decision-v7 replay 4,000 | 1 epoch, lr 2e-5 |
| v0.2 | Japanese and agent routing | Japanese decisions from JGLUE (JNLI, JCommonsenseQA, JSTS; 6,000), agent-factory and model-routing questions (2,745), decision-v7 replay 3,000 | 1 epoch, lr 5e-5 |
| v0.3 | Pull-request labeling | 3,500 English and 768 Japanese pull requests labeled with change type, blast radius and severity (closed PRs of NousResearch/hermes-agent, MIT; Japanese copies machine-translated; severity follows the labeling job's rules for docs-only and dependency PRs), replay: decision-v7 1,500, JGLUE 500, routing 300 | 1 epoch, lr 5e-5, documents up to 1,536 tokens |
| v0.4 | Pull-request labeling, round 2 | the 4,063 English PRs v0.3 did not see, its 2,396 rare-label PRs (blast radius, P0, P1, P4) again, 2,185 more PRs with a blast-radius label (blast and severity questions only; all older than every evaluation PR), the 768 Japanese PRs again; replay: decision-v7 1,500, JGLUE 500, routing 300 (11,540 records after dropping 172 too long) | from v0.3, 1 epoch, lr 5e-5, documents up to 1,536 tokens, 1,443 steps on one L4 |
Calibration: a single temperature, T = 1.59, fitted on pooled held-out rows (decision-v7 calibration split and PR-labeling development set; calibration error 0.066 before, 0.019 after, out of fold), stored in head.pt.
Results
Every row scored through the same HTTP client (scripts/compare_systemone.py) on the MLX 8-bit builds, held-out data only:
| Set | v0.3 | v0.4 |
|---|---|---|
| PR labels, English (953 PRs): change type / severity / blast radius | 89% / 70% / 54% | 88% / 78% / 51% |
| PR labels, English, all labels | 78% | 81% |
| PR labels, English, labels where p ≥ 0.7 | 93% right on 57% of labels | 92% right on 65% of labels |
| PR labels, Japanese (92 PRs) | 70% | 74% |
| 17 real open PRs from the owner's other repositories (labels drafted by Claude) | 61% | 69% |
| JGLUE development (500, Japanese) | 85% | 82% |
| Held-out transfer-v4 development (764, sources never trained on) | 71% | 69% |
| Agent factory development (640 questions) | 92% | 91% |
| Model routing: generic development (270) / hand-labelled (45) | 99% / 100% | 97% / 100% |
Blast radius: v0.3 answered "broad" for 81 of 139 test PRs; v0.4 spreads its answers (33 broad, 67 moderate, 39 contained) and finds more contained (20 of 42, was 14) and moderate (28 of 44, was 20) PRs, fewer broad ones (23 of 44, was 41). Its accuracy is about the same; it is less biased.
The questions it was trained for
PR labeling works best with these questions, worded exactly so, over a document of the form title / author / stats / body / files (one - status path +added/-deleted line per file):
type(choice): "Primary change type from files and body, not the title prefix." Options: type/bug, type/docs, type/feature, type/perf, type/refactor, type/security, type/test.blast(choice): "How far a mistake in this PR spreads in production." Options: review:blast-contained (one module), -moderate (one subsystem), -broad (shared helper or config), -massive (auth, permissions, or all paths).sev(choice): "How serious the problem this PR addresses is — not the risk of merging the diff as-is." Options P0 to P4.
The PR-labeler recipe has the exact wording, the document builder and the training and scoring scripts. Other questions about other documents work as in v0.2.
Run it
The checkpoint runs on any device; no retraining per device:
pip install "d1a[serve] @ git+https://github.com/jonpol01/d1a"
python -m d1a.serve --run JohnP1/d1a-e4b@v0.4 --port 8009 # NVIDIA GPU, CPU, or Apple GPU (PyTorch)
python -m d1a.serve --run JohnP1/d1a-e4b-mlx-q8@v0.4 --port 8009 # Apple Silicon: the 8-bit MLX build (~6 GB)
or in-process: from d1a import D1A; D1A.load("JohnP1/d1a-e4b@v0.4").decide(state, questions).
Limits
- Blast radius is still the weakest label (51% on the English test) and has never picked "massive" (0 of 9; only 87 such PRs exist in the source project). A "needs a human" rule should not rest on it alone.
- Severity is still weakest at the urgent end: on the English test it finds 13 of 38 P0 and 17 of 47 P1 PRs (v0.3: 16 and 15), against 505 of 564 P3 and 75 of 92 P4.
- The earlier skills are 1–3 points below v0.3 (table above). For general decisions, Japanese JGLUE-style questions or routing alone, v0.3 is the better choice.
- The PR data comes from one large open-source project; label conventions follow its maintainers except where the labeling job's rules override them.
- Questions are in English; the documents can be English or Japanese.
Versions
| Tag | What it adds |
|---|---|
v0.1 |
general decisions (decision-v7, 2 epochs, calibrated) |
v0.2 |
Kev's later skills, Japanese, agent routing |
v0.3 |
pull-request labeling, English and Japanese |
v0.4 |
pull-request labeling round 2: severity, Japanese, less biased blast radius |
Older tag names (v0.1-2epoch, v0.1.1-2epoch-calibrated, v0.2-hybrid) still work.
License and data
Apache-2.0. Base model: Gemma 4 by Google (Apache-2.0). Code: github.com/jonpol01/d1a, built on Kev by Jared Palmer (Apache-2.0). Japanese decision data derived from JGLUE by Yahoo Japan Corporation and Waseda University (CC BY-SA 4.0). Pull-request data from NousResearch/hermes-agent (MIT); the labeled PR dataset itself is private.
- Downloads last month
- 39