D1A-E2B 路 MLX q8 (per-layer embeddings 4-bit)

The D1A-E2B v0.2 decision model (two epochs, JohnP1/d1a-e2b) converted for Apple Silicon: the trained adapter merged into Gemma 4 E2B, linear layers and token embeddings at 8 bits, the large per-layer embedding table at 4 bits, and the pointer head in fp32 (head.safetensors). Typed questions in, a calibrated probability for every option out, one forward pass. Same System One API as Jev.

This MLX build PyTorch bf16
Memory after load / peak (60 questions) 4.2 GB / 6.4 GB ~10-12 GB / ~15 GB
Load time 1.4 s 16-20 s
6-question request (M1 Max) ~480 ms cold, ~400 ms warm similar

Parity against the PyTorch reference path (bf16, the precision v0.2 was trained and evaluated in), measured on 274 questions (decision-v7 development records, the playground presets, long states): mean |dp| 0.009, p95 0.038, max 0.17, 5 changed answers (1.8%, the same rate as the v0.1 build), accuracy 0.824 vs 0.828, ECE 0.066 vs 0.072; the playground presets change no answer. Tag v0.1-1epoch keeps the build of v0.1. A pure 4-bit build failed the same gate (8.4% changed answers) and is not published.

Run it

pip install "d1a[serve] @ git+https://github.com/jonpol01/d1a@mlx-gemma4"
python -m d1a.serve --run JohnP1/d1a-e2b-mlx-q8 --port 8009

Apple Silicon only (MLX). Playground: https://github.com/jonpol01/d1a-playground

License

Apache-2.0. Base model: Gemma 4 by Google (Apache-2.0). Code: github.com/jonpol01/d1a, built on Kev by Jared Palmer (Apache-2.0).

Downloads last month
-
Safetensors
Model size
5B params
Tensor type
U32
路
BF16
路
MLX
Hardware compatibility
Log In to add your hardware

8-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 馃檵 Ask for provider support

Model tree for JohnP1/d1a-e2b-mlx-q8

Quantized
(1)
this model