- Gemma 4 E2B β Core AI (.aimodel)
- Provenance
- Contents
- Requirements
- Performance
- Status
- License
stable-smoke/β truncated proving assetstable/β shape-stable decode contractring-smoke/β truncated proving assetring/β shape-stable decode, ring sliding cachesring2/β the same graph at three KV capacitiesring2-smoke/β truncated proving assetring3/β per-layer-type SDPA formring3-smoke/β truncated proving asset
- Provenance
Gemma 4 E2B β Core AI (.aimodel)
google/gemma-4-E2B-it-qat-q4_0-unquantized converted to Core AI .aimodel bundles for
Apple silicon by visible-cx. These are derivative
artifacts: Google's QAT-trained weights rounded onto the int4 grid they were trained for
and re-expressed as a Core AI graph. They load through Core AI on macOS and are not usable
by PyTorch, GGUF or MLX.
Gemma 4 E2B uses Per-Layer Embeddings, so these bundles take a large embedding gather
table as a static graph input rather than carrying it in the graph. That table ships in
ios-frontend/ and the bundles do not load without it; a missing table produces a bare
input-arity error naming ple_table/ple_scale.
β οΈ Known issue β memory growth on long generations. The Core AI runtime caches a graph specialization per input-shape signature, and this export makes the sequence length load-bearing: every generated token adds one token to
position_idsand thus mints a new signature, retaining roughly 80 MB of GPU allocations per generated token (both engines, guided and unguided alike, at equal context). Workloads whose prompt/output lengths repeat β e.g. short constrained extraction rows β stay bounded, because previously seen lengths replay from cache at no cost. Long free-form generations grow without bound until the process exits; no in-process mitigation exists (token caps only postpone it, and shape-length reuse is the only stable state). A re-export onto a shape-stable contract that removes this growth is available in this repository understable/β see its section below. It requires updated host support (in progress); on current hosts, use the bundles above for guided generation only.
Provenance
| Base checkpoint | google/gemma-4-E2B-it-qat-q4_0-unquantized (ungated) |
| Zoo recipe | gemma-4-e2b, status = "verified" β int4lin --tbl |
| Recipe (pf64) | export_gemma4_pf_pipelined.py --pf 64 with --tbl and --raw-dir pointed at the gather table below |
| Toolchain base | apple/coreai-models @ b1cb71b8522d99408059fa0b98b8742171bcb0b8 + the coreai-model-zoo python overlay |
| Toolchain | coreai-torch 0.4.1, coreai-core 1.0.0b2, coreai-opt 0.2.1, torch 2.9.0 |
| Producer fingerprint | coreai-core 1.0.0b2 on every inner asset metadata.json |
| Weight format | int4, per-block-32 (int4lin, symmetric-with-clipping) β the ggml q4_0 grid the QAT checkpoint was trained on |
| Vocab | 262,144 |
| Export functions | main (S=1 decode) and, in _pf64 bundles, prefill (S=64 chunked prefill) |
"QAT-unquantized" means QAT-trained, stored at full width; the int4 rounding happens at export, onto the grid training already targeted.
_tbl = the PLE gather table is bound as a static graph input. _pf64 = a second
entrypoint, prefill, with a static query width of 64
(function_map: {"main": ["main", "prefill"]}).
Contents
Bundles
| Path | Bytes | Context | Functions |
|---|---|---|---|
gpu-pipelined/gemma4_e2b_qat_decode_int4lin_tbl_pf64 |
2,122,089,973 | 4096 | main + prefill |
w4a8/gemma4_e2b_qat_decode_int4lin_a8_tbl_pf64 |
2,122,679,604 | 16384 | main + prefill |
Each bundle folder holds <name>.aimodel/ (main.mlirb β 2.09 GB, main.hash, asset
metadata.json), a bundle-level metadata.json, and tokenizer/ (tokenizer.json
32,169,626 B, tokenizer_config.json, generation_config.json, chat_template.jinja
18,569 B). The w4a8 folder additionally ships its calibration_corpus.jsonl (35,045 B).
Stop token: every bundle declares eos_token = "<turn|>" (id 106), which is the
turn terminator Gemma 4 emits. generation_config.json independently lists
eos_token_id: [1, 106, 50]. A host that stops on the raw upstream <eos> instead will
overrun every reply.
The PLE gather-table sidecar β required, not optional
| Path | Files | Bytes |
|---|---|---|
ios-frontend/gemma4_qat_gather_raw/ |
7 | 2,808,612,071 |
| File | Bytes |
|---|---|
embed_per_layer.i8 |
2,348,810,240 |
embed_tokens.i8 |
402,653,184 |
proj.f32 |
55,050,240 |
embed_per_layer.scale.f32 |
1,048,576 |
embed_tokens.scale.f32 |
1,048,576 |
proj_norm.f32 |
1,024 |
meta.json |
231 |
meta.json records the shape and the dequant convention:
V 262144, D 1536, PLD 8960, L 35, ld 256, embed_scale_pl 16.0 (= β256, which is what the
exporter assumes).
Every _tbl bundle binds ple_table (from embed_per_layer.i8, 2.35 GB) and ple_scale
(from embed_per_layer.scale.f32) as static graph inputs.
Requirements
Apple silicon Mac, Core AI runtime.
Engine contract: 4 inputs β
input_ids,position_ids, plus staticple_tableandple_scale. Two engines accept that:- Pipelined engine β binds the statics zero-copy over the caller's buffer, but does not support logits, so no grammar-constrained decoding.
- Sequential engine β the only logits-capable engine, and therefore the only path for
guided decoding. It accepts
>= 2inputs and binds everything beyondinput_ids/position_idsfromEngineOptions.staticInputBuffers; an extra input with no buffer is rejected by name. It pays a one-time copy of every constant input at load β 2.19 GB for E2B, because its submission path otherwise materialises a foreign buffer-backed view on every forward pass. Binding the table as a raw view instead costs ~0.30 s per pass (3.4 tok/s); the load-time copy takes the same bundle to 8.6 s/row.
States:
keyCache/valueCacheFloat16, 15 Γ 1 Γ 1 Γ ? Γ 512. Dynamic sequence dim βGrowingKVCache(initial 256, doubling), not a static allocation at the manifest maximum.KV cost: 30,720 bytes per token (fp16) β 126 MB at 4096, 252 MB at 8192, 503 MB at
Memory envelope, 16 GB Mac (Metal
recommendedMaxWorkingSetSizeβ 10.7 GB):Weights ( main.mlirb)2.09 GB PLE table bound as static input 2.35 GB Base 4.44 GB Affordable context, fp16 KV ~204k tokens β 131,072, the model cap, binds first KV is not the ceiling at this tier; the PLE table is. Measured peak footprint on a guided run: 8.26 GB, max RSS 5.40 GB. Minimum practical machine memory: 16 GB, for guided decoding only. Free-form generation does not fit this envelope on any machine β see the known issue above.
Sliding-window note: E2B interleaves 12 sliding and 3 full attention layers across 35 layers collapsed to 15 KV slots. The export models the window in the mask, not in memory β sliding layers occupy full-length slots and zero-pad head_dim 256β512, so 80% of KV bytes sit in windowed slots and 40% is dead padding.
The bundle manifest declares
runtime_env COREAI_CHUNK_THRESHOLD=1. Both engines read the static query width off the graph.
Performance
Measured on a 16 GB Apple silicon Mac (M2 Pro, macOS 27 beta): 10-sample harness,
guided JSON-constrained decoding against a fixed schema, greedy, sequential engine with the
PLE tables bound as static inputs, reset() between samples, 128-token cap. Load excluded
from s/row; sample 1 excluded as a cache-warm outlier.
_tbl_pf64 |
|
|---|---|
| Load | 12.1 s |
| Guided JSON parse | 10/10 |
| Enum-clean | 10/10 |
| s/row (long samples) | 6.12 |
| s/row (short samples) | 4.23 |
| Decode | 22.7β32.7 tok/s |
| TTFT | 0.59β4.20 s |
| Peak footprint | 8.26 GB |
| Max RSS | 5.40 GB |
| Outcome | completed all ten samples |
The S=64 prefill function carries the time-to-first-token. Measured head to head on
a 942-token prompt against a decode-only export of the same weights, exporting prefill
moved TTFT 35.74 s β 4.65 s and s/row 45.04 β 8.83, a 7.7Γ. Decode is untouched β the
prefill function only serves full-width chunks. The decode-only bundle is no longer
published; every bundle in this repo now carries prefill.
Enum conformance is the grammar's doing. Unguided, the model emits an off-schema enum value on essentially every sample. Guided, all ten are correct, because an off-enum token is unsamplable. Decode throughput is essentially unchanged by the constraint; the ~25% extra cost of a guided sample is the sequential engine's step-synchronous prefill.
For reference, upstream measured E2B at 77.0/87.1 tok/s on an M4 Max, and an unguided pipelined run on the 16 GB machine reached 44.3 tok/s. The 22.7β32.7 tok/s above is the guided, sequential-engine figure.
Unguided workload β memory-capped
Free-form generation from a fixed prompt. Every generation length below is a memory cap, not a model stop β see the known issue.
| depth | prompt tokens | TTFT | decode | generated | wall | peak footprint |
|---|---|---|---|---|---|---|
| 3.4k | 3314 | 12.21 s | 7.2 tok/s | 128 (capped) | 30.4 s | 19.97 GB |
| 8k | 7972 | 54.01 s | 2.1 tok/s | 64 (capped) | 84.3 s | 36.71 GB |
Output quality up to the cap is sound; the limit is memory, not capability.
Needle-in-haystack recall: 3/3 verbatim at 8k, within a 64-token cap (three distinctive facts planted at 10/50/90% of the filler, strict scoring). 15k was not attempted: 8k already cost 36.71 GB of footprint.
Status
| Artifact | Status |
|---|---|
gpu-pipelined/gemma4_e2b_qat_decode_int4lin_tbl_pf64 (ctx 4096) |
QUALIFIED FOR GUIDED GENERATION β measured: 10/10 parse, 10/10 enum-clean, 6.12/4.23 s/row, 22.7β32.7 tok/s, 8.26 GB peak footprint. Not qualified for free-form generation on the current reference engine implementation. |
w4a8/gemma4_e2b_qat_decode_int4lin_a8_tbl_pf64 |
EXPERIMENTAL β built, unmeasured. Gate: a Mac-side oracle/parity check plus a device benchmark. |
stable/β¦_tbl_pf64_stable_c16384 |
EXPERIMENTAL β shape-stable decode contract; removes the per-token memory growth by design; host support exists; measured E2B: memory fix confirmed but ~1 tok/s decode at capacity 16384 β awaiting a smaller-capacity export before any use. |
stable-smoke/β¦_stable_c16384_l5 |
NOT A MODEL β 5-layer truncated proving asset for host development; produces low-quality text by design. |
About the w4a8 bundle
Same QAT int4 per-block-32 weights and the same graph as
gpu-pipelined/gemma4_e2b_qat_decode_int4lin_tbl_pf64; the difference is an int8
quantize/dequantize pair on the inputs of every F.linear, calibrated on 128 synthetic
samples (the corpus ships in the folder). It was built as a prefill/TTFT lever.
The caveat comes from the compression toolchain's own documentation: the runtime fast path
is W_INT8 Γ A_INT8, and a float weight path executes in floating point regardless of
activation quantization. These weights use int4 scale-multiply dequant, not an int8 LUT, so
the expected TTFT gain may be zero or negative. This bundle exists to be measured, not
because a win is predicted.
No numerics gate has been run on any bundle in this repo. The 10/10 results are behavioural (parse rate, enum conformance, clean stop); a decode oracle against an fp32 reference has not been run.
License
The upstream QAT checkpoint carries Apache-2.0 metadata and is ungated, but it is a Gemma
model and the Gemma terms apply downstream, which is why this repo declares
license: gemma. Use is subject to the
Gemma Terms of Use, the
Gemma Prohibited Use Policy and the
Gemma 4 license. Those obligations
travel with any redistribution of these bundles, including the gather-table sidecar, which
is derived from the same weights. The contribution here is the conversion, not the weights.
stable-smoke/ β truncated proving asset
β οΈ EXPERIMENTAL β not a usable model. A 5-layer truncation of the E2B decoder, published only so the shape-stable engine contract below can be developed against a small download. It carries real weights for the layers it keeps and produces low-quality text; do not evaluate quality from it.
| Bundle | Bytes | Context | Functions | Producer | Created |
|---|---|---|---|---|---|
stable-smoke/gpu-pipelined/gemma4_e2b_qat_decode_int4lin_tbl_pf64_stable_c16384_l5 |
1,171,287,831 | 16384 | main+prefill | coreai-core 1.0.0b2 | 20260818T003645Z |
stable/ β shape-stable decode contract
β οΈ EXPERIMENTAL β built, not yet measured on a Mac. A re-export of the same weights and the same quantization onto a decode/prefill contract in which no input shape moves between steps, which removes the per-generated-token memory growth described in the known issue above. It requires a host that feeds
position_idsas the absolute positions of the S new tokens only (host support in progress); a host that feeds the growing0..Nprefix will write the KV cache at the wrong offset. Measured on host support (16 GB M2 Pro): the memory defect is confirmed fixed β 6.53 GB flat peak across a 664-token free-form generation (+0.011 MB/token, versus ~81 MB/token and a killed process on the default bundle) and guided output byte-identical to the default bundle β but decode runs at ~1.0 tok/s versus ~27β31, because every step reads the full 16,384-slot cache. Not usable as shipped; a smaller-capacity export would trade window for speed. Treat as a working proof of the contract, not a deployable bundle.
| Bundle | Bytes | Context | Functions | Producer | Created |
|---|---|---|---|---|---|
stable/gpu-pipelined/gemma4_e2b_qat_decode_int4lin_tbl_pf64_stable_c16384 |
2,122,071,656 | 16384 | main+prefill | coreai-core 1.0.0b2 | 20260818T003828Z |
Contract. Both entrypoints take four inputs and two states, all statically shaped:
main IN input_ids Int32 1x1 | position_ids Int32 1x1 | ple_table Int8 V x (L*ld) | ple_scale Float32 V
prefill IN input_ids Int32 1x64 | position_ids Int32 1x64 | (same statics)
ST keyCache / valueCache Float16 [slots, 1, n_kv, 16384, 512]
OUT logits Float16 1 x S x 262144
position_ids carries the absolute position of each of the S tokens in the call, and
position_ids[0,0] is also the cache slot the K/V for those tokens is written at. The
attention mask (causality, unwritten slots, and the 512-token sliding window) is derived
inside the graph from those positions, so there is no mask input to bind. The KV state's
sequence extent is a literal 16,384 rather than a dynamic dimension, so a host that
resolves its cache strategy from the state descriptor will allocate the full cache up
front (0.50 GB for E2B, 1.61 GB for E4B) instead of growing it β that fixed cost is the
trade for the unbounded growth it replaces.
The context ceiling here is encoded in the graph: a different window needs a different export, not a manifest edit.
Verification. Before conversion, the stable graph was checked against the shipped graph in eager torch on the same weights: greedy argmax identical on all tested decode steps (including positions past the 512-token sliding window, with 384 cache slots unwritten-and-masked), worst logits cosine 0.99998 (fp16), and the prefill path bit-identical in fp32. On-device behavior is unmeasured until host support lands.
ring-smoke/ β truncated proving asset
β οΈ EXPERIMENTAL β not a usable model. A 5-layer truncation of the E2B decoder, published only so the shape-stable engine contract can be developed against a small download. It carries real weights for the layers it keeps and produces low-quality text; do not evaluate quality from it.
| Bundle | Bytes | Context | Functions | Producer | Created |
|---|---|---|---|---|---|
ring-smoke/gpu-pipelined/gemma4_e2b_qat_decode_int4lin_tbl_pf64_ring_c16384_l5 |
1,171,301,346 | 16384 | main+prefill | coreai-core 1.0.0b2 | 20260818T085021Z |
ring/ β shape-stable decode, ring sliding caches
β οΈ EXPERIMENTAL β built and gated in torch, on-device measurements pending. A re-export of the same weights and the same quantization as
stable/, onto the same host contract, with the KV layout reworked. No throughput or footprint numbers have been taken on a Mac yet; the structural facts below are read off the converted program, not inferred.
| Bundle | Bytes | Context | Functions | Producer | Created |
|---|---|---|---|---|---|
ring/gpu-pipelined/gemma4_e2b_qat_decode_int4lin_tbl_pf64_ring_c16384 |
2,122,101,099 | 16384 | main+prefill | coreai-core 1.0.0b2 | 20260818T085141Z |
What this is. stable/ removed the per-generated-token memory growth by making every
input shape a literal, and it worked β but it gave every layer a full 16,384-slot KV cache
and read all of it every step. ring/ keeps that contract byte-for-byte and changes what
sits behind it:
- sliding layers read a ring buffer, 576 slots (the model's own 512-token sliding
window plus one 64-token prefill chunk) instead of 16,384. Ring slot
rholds the newest written position withpos % 576 == r, and the mask that selects the written, causal, in-window slots is derived inside the graph from the absolute positions. - the grouped-query head expansion is folded into the query rather than materialised.
The stock lowering broadcasts an index tensor to the full
[1, heads, K, head_dim]and gathers K and V through it; reshaping the query so the head counts already match produces the same dot products with no index tensor and no gathered copies. - both regions are packed into the same two states, so the host still binds exactly two KV states, positionally, at whatever literal extent the descriptor declares.
Contract β unchanged from stable/. Both entrypoints take four inputs and two states,
all statically shaped:
main IN input_ids Int32 1x1 | position_ids Int32 1x1 | ple_table Int8 V x (L*ld) | ple_scale Float32 V
prefill IN input_ids Int32 1x64 | position_ids Int32 1x64 | (same statics)
ST keyCache / valueCache Float16 [1, 1, n_kv, TOTAL, 512]
OUT logits Float16 1 x S x 262144
position_ids carries the absolute position of each of the S tokens in the call, and
position_ids[0,0] is the write position. The mask is derived in the graph, so there is
no mask input to bind. The KV states have literal extents, so a host that resolves its
cache strategy from the state descriptor allocates them up front rather than growing them.
Structural facts (dumped from the converted program):
| E2B | E4B | |
|---|---|---|
| KV state shape | [1, 1, 1, 56064, 512] x2 |
[1, 1, 2, 77056, 512] x2 |
| KV bytes, both states | 114.8 MB | 315.6 MB |
same for stable/ |
503.3 MB | 1.61 GB |
| cache slots read per decode step | 56,064 | 77,056 |
same for stable/ |
245,760 | 393,216 |
| sliding / full layers | 12 x 576 + 3 x 16,384 | 20 x 576 + 4 x 16,384 |
| dynamic dimensions anywhere | none | none |
One new host precondition. A prefill call's first position must be a multiple of 64.
The ring write is a fixed-width store at p0 % 576, and 576 is nine 64-token chunks, so a
64-aligned chunk can never straddle the wrap; an unaligned one would write past the end of
its own ring region. Feeding whole 64-token chunks from a 64-aligned position β and the
remainder one token at a time through main β was already what the contract asked for.
stable/ tolerated an unaligned chunk; ring/ does not.
Unlike the ctx*/ folders, the context ceiling is encoded in the graph: a different
window needs a different export, not a manifest edit.
ring2/ β the same graph at three KV capacities
EXPERIMENTAL β diagnostic assets, not products. These are the
ring/recipe, byte for byte, exported at a smaller full-attention capacity. They exist to measure how the per-step cost of the shape-stable graph depends on that capacity. No throughput numbers have been taken on device for either of them. Their context ceilings (2,048 and 8,192) are smaller than the 16,384 ofring/and are encoded in the graph, so neither is a drop-in replacement for anything already published here.
| Bundle | Context | KV state shape | KV bytes, both states | Sliding / full layers |
|---|---|---|---|---|
ring2/gpu-pipelined/gemma4_e2b_qat_decode_int4lin_tbl_pf64_ring_c2048 |
2,048 | [1, 1, 1, 13056, 512] x2 |
26.7 MB | 12 x 576 + 3 x 2,048 |
ring2/gpu-pipelined/gemma4_e2b_qat_decode_int4lin_tbl_pf64_ring_c8192 |
8,192 | [1, 1, 1, 31488, 512] x2 |
64.5 MB | 12 x 576 + 3 x 8,192 |
ring/gpu-pipelined/..._ring_c16384 (for comparison) |
16,384 | [1, 1, 1, 56064, 512] x2 |
114.8 MB | 12 x 576 + 3 x 16,384 |
Contract: identical to ring/. Same four inputs in the same order, same two states in
the same positional order, position_ids as the absolute positions of the S tokens in the
call, the mask derived in the graph, the same 64-alignment precondition on prefill. The
only thing that differs is the literal extent of the full-attention region, and a host that
reads its cache strategy from the state descriptor needs no change to drive them.
Equivalence. Both were gated in eager torch against the shipped graph before conversion, at their own capacity: 1,600 prompt tokens as 25 chunks of 64 then 32 greedy decode steps, final position 1,631, so the sliding ring wraps roughly three times. fp32, 57 comparisons per run, 0 argmax mismatches and worst cosine 0.999999999987 for both. No dynamic dimension appears in either entrypoint of either bundle.
ring2-smoke/ β truncated proving asset
EXPERIMENTAL β not a usable model. A 5-layer truncation of the E2B decoder at the 2,048 capacity (
[1, 1, 1, 4352, 512]states, 4 x 576 + 1 x 2,048), published only so host-side work can be developed against the contract over a smaller download. It carries real weights for the layers it keeps and produces low-quality text; do not evaluate quality from it.
ring3/ β per-layer-type SDPA form
EXPERIMENTAL β gated in torch, no on-device numbers yet. Same weights, same quantization, same host contract as
ring/. One thing changes inside the graph: the grouped-query expansion is folded into the query only in the full-attention layers, and the sliding layers keep the stock gathered form.
Why the split. Folding the GQA expansion into the query removes the gathered copies and
the broadcast index tensor, but it also turns G independent per-head matmuls into a single
batch-1 matmul with G times the rows. That trade is worth taking when the key length is long
(the full-attention layers read the whole capacity) and not worth taking when it is short
(the sliding layers read a fixed 576-slot ring, where the operation is short enough that
losing the per-head parallelism costs more than the bytes it saves). ring/ applied the fold
everywhere; ring3/ applies it only where the key length is long.
Both forms compute the same dot products in the same order. The equality gate below is run
against the shipped graph, not against ring/.
Contract: unchanged. Same four inputs in the same order, same two KV states in the same
positional order, position_ids carrying the absolute position of each of the S tokens in
the call, the mask derived inside the graph, the same requirement that a prefill call start
at a multiple of 64. A host driving ring/ drives these with no change.
Equality. Gated in eager torch against the shipped graph before conversion: 1,600 prompt tokens as 25 chunks of 64 then 32 greedy decode steps, final position 1,631, so the sliding ring wraps about three times. fp32, 57 comparisons per run, 0 argmax mismatches on every bundle below. No dynamic dimension appears in either entrypoint of any of them.
| Bundle | Context | KV state shape | KV bytes, both states | Sliding / full slots |
|---|---|---|---|---|
ring3/gpu-pipelined/gemma4_e2b_qat_decode_int4lin_tbl_pf64_ring_c4096_gqafull |
4,096 | [1, 1, 1, 19200, 512] x2 |
39.3 MB | 12 x 576 + 3 x 4,096 |
ring3/gpu-pipelined/gemma4_e2b_qat_decode_int4lin_tbl_pf64_ring_c2048_gqafull |
2,048 | [1, 1, 1, 13056, 512] x2 |
26.7 MB | 12 x 576 + 3 x 2,048 |
E2B runs 28 sliding and 7 full attention layers; in these bundles the 28 keep the gathered form and the 7 use the folded one.
ring3-smoke/ β truncated proving asset
EXPERIMENTAL β not a usable model. A 5-layer truncation of the E2B decoder at the 4,096 capacity (
[1, 1, 1, 6400, 512]states, 4 x 576 + 1 x 4,096), published only so host-side work can be developed against the contract over a smaller download. It carries real weights for the layers it keeps and produces low-quality text; do not evaluate quality from it.
Model tree for visible-cx/gemma-4-E2B-CoreAI
Base model
google/gemma-4-E2B