gemma-4-e4b-it β€” q8_0 qcache (candle ignite-perf format)

Quantized serving set for google/gemma-4-e4b-it (3n-class MatFormer architecture: per-layer embeddings, altup/laurel blocks), produced and consumed by the ignite-perf candle fork (branch ignite-perf).

Files

file contents
qcache-q8_0.v1.gguf all quantized linear projections (q8_0)
aux-tensors.safetensors every non-quantized tensor β€” for the 3n family this includes the large per-layer embedding tables, altup/laurel weights, norms (hence the size)
config.json, tokenizer* upstream configs

qcache + aux is a complete runnable set β€” the original bf16 shards are not required.

Loading (ignite-catalog exp1 executor)

{
  "asset_prefix": "<this repo's local dir>",
  "model_family": "gemma4",
  "quantize": "q8_0",
  "model_files": ["aux-tensors.safetensors"]
}

Notes

  • Quantization is deterministic (byte-identical across regenerations).
  • Use of this model is subject to the upstream Gemma license terms.
Downloads last month
4
GGUF
Model size
5B params
Architecture
Hardware compatibility
Log In to add your hardware

8-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support