gemma-4-e4b-it β q8_0 qcache (candle ignite-perf format)
Quantized serving set for google/gemma-4-e4b-it (3n-class MatFormer
architecture: per-layer embeddings, altup/laurel blocks), produced and
consumed by the ignite-perf candle fork
(branch ignite-perf).
Files
| file | contents |
|---|---|
qcache-q8_0.v1.gguf |
all quantized linear projections (q8_0) |
aux-tensors.safetensors |
every non-quantized tensor β for the 3n family this includes the large per-layer embedding tables, altup/laurel weights, norms (hence the size) |
config.json, tokenizer* |
upstream configs |
qcache + aux is a complete runnable set β the original bf16 shards are
not required.
Loading (ignite-catalog exp1 executor)
{
"asset_prefix": "<this repo's local dir>",
"model_family": "gemma4",
"quantize": "q8_0",
"model_files": ["aux-tensors.safetensors"]
}
Notes
- Quantization is deterministic (byte-identical across regenerations).
- Use of this model is subject to the upstream Gemma license terms.
- Downloads last month
- 4
Hardware compatibility
Log In to add your hardware
8-bit
Inference Providers NEW
This model isn't deployed by any Inference Provider. π Ask for provider support