gemma-4-E2B-it-q4.apml — Apertura model bundle

Gemma 4 E2B (elastic 2B) instruction-tuned, quantized to the Apertura .apml bundle format: MLX-affine 4-bit (group size 64, 8-bit embeddings), tokenizer.json + chat_template.jinja included, runtime mlx, architecture gemma4. Post-training quantization of the bf16 release (this family has no QAT variant).

Consumed by Apertura — an Objective-C++ transformer engine + macOS chat app on MLX. Export gate: --verify-bundle (bundle reload == in-memory quantization, argmax-identical).

Download

hf download apocryphx/gemma-4-E2B-it-q4-apml --local-dir gemma-4-E2B-it-q4.apml

License

Gemma is provided under and subject to the Gemma Terms of Use. By downloading this model you agree to those terms, including the Gemma Prohibited Use Policy. This repository redistributes a quantized Model Derivative of google/gemma-4-E2B-it; the same terms and use restrictions apply to it and to any further derivatives.

Downloads last month
22
MLX
Hardware compatibility
Log In to add your hardware

Quantized

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for apocryphx/gemma-4-E2B-it-q4-apml

Finetuned
(342)
this model