Access Gemma on Hugging Face

This repository is publicly accessible, but you have to accept the conditions to access its files and content.

To access Gemma on Hugging Face, you must review and agree to Google's Gemma Terms of Use and Prohibited Use Policy. Requests are processed immediately.

Log in or Sign Up to review the conditions and access this model content.

Gemma 4 12B IT — 4-bit MLX

A 4-bit quantization of google/gemma-4-12B-it for MLX on Apple Silicon.

Produced for Triad, a local three-model token-fusion experiment. These are the exact weights that project was built and validated against, published so its results can be reproduced.

Size on disk: 6.3 GB (down from ~24.8 GB at bf16)

⚠️ License notice — read before using

Gemma is not released under an open-source license. It is governed by the Gemma Terms of Use and the Gemma Prohibited Use Policy, which impose restrictions that Apache-2.0 and similar licenses do not.

By downloading or using this model you agree to those terms, exactly as if you had obtained it from Google directly. Quantization creates a derivative — it does not relicense anything. If you redistribute this model or anything derived from it, you are responsible for passing these terms and restrictions on to your recipients.

Some copies of Gemma quantizations circulating on the Hub are mislabelled apache-2.0, including the upstream metadata this conversion inherited. That label was incorrect and has been corrected here. Do not rely on a license tag alone — read the terms.

Requirements

  • Apple Silicon
  • A build of mlx-lm that supports gemma4_unifiedthe PyPI release does not

⚠️ pip install mlx-lm is not sufficient for this model

This model's config declares model_type: gemma4_unified. mlx-lm resolves that through its MODEL_REMAPPING table (gemma4_unified → gemma4), but that entry is not present in the PyPI release — verified against the mlx-lm 0.31.3 wheel, whose MODEL_REMAPPING has no gemma4_unified key and which ships no models/gemma4_unified.py. Installing from PyPI and calling load() on this repo will fail to resolve the architecture.

Install from the git commit this model was built and verified against:

pip install "mlx-lm @ git+https://github.com/ml-explore/mlx-lm.git@cf10f962b7a20e63a6df43dbf0faf06070153d40"

Once a PyPI release includes gemma4_unified in MODEL_REMAPPING, plain pip install mlx-lm will work and this note can be ignored. Check before assuming either way.

No packages beyond mlx-lm are required at load time. (mlx-optiq is used during conversion — see Quantization provenance — but is not needed to load this repo: verified by blocking the optiq import and loading successfully.)

Usage

from mlx_lm import generate, load

model, tokenizer = load("Micklavin/gemma-4-12B-it-4bit")

prompt = tokenizer.apply_chat_template(
    [{"role": "user", "content": "What is the chemical symbol for gold?"}],
    tokenize=False,
    add_generation_prompt=True,
)
print(generate(model, tokenizer, prompt=prompt, max_tokens=256))

Thinking mode

Gemma 4's chat template branches on enable_thinking and emits a <|channel>thought block. Triad disables it so all ensemble members answer in the same phase:

prompt = tokenizer.apply_chat_template(
    messages, tokenize=False, add_generation_prompt=True, enable_thinking=False
)

Note that with thinking disabled this template still opens a thought channel where other models' templates close theirs. If you filter model output, account for that.

Quantization provenance

Converted with mlx_lm.convert:

convert(
    hf_path="google/gemma-4-12B-it",
    mlx_path="models/gemma-4bit",
    quantize=True,
    q_bits=4,
    q_group_size=64,
    dtype="bfloat16",
)

Resulting config: {"group_size": 64, "bits": 4, "mode": "affine"}, model_type: gemma4_unified. Per mlx-lm's remapping comment, this is the encoder-free multimodal variant with vision and audio weights stripped by sanitize() — so this is a text-only conversion.

Component Version
mlx 0.32.0
mlx-lm 0.31.3
transformers 5.12.1

The reproduction script is scripts/quantize_models.py.

Evaluation

None beyond a smoke test. No benchmark comparison against the bf16 original was run, so the quantization's quality cost is unmeasured. Treat it as an untested 4-bit conversion rather than a validated one.

License

Gemma Terms of Use. See the notice above — this is a custom license with use restrictions, not an open-source one.

Downloads last month
8
Safetensors
Model size
2B params
Tensor type
BF16
·
U32
·
MLX
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Micklavin/gemma-4-12B-it-4bit

Quantized
(299)
this model