Reflex-1-4B-bnb-4bit

Pre-quantized 4-bit (bitsandbytes NF4, double quantization) copy of Google Gemma 4 E4B for use as the base model of Reflex-1. Download is 9.3 GB instead of about 16 GB, and no quantization is needed at load time.

The vision/audio towers, embeddings and LM head are kept in bfloat16; the transformer blocks are NF4.

Variants

All Reflex-1 repositories: Reflex-1 collection

Repository What it is Download GPU memory
matu79go/Reflex-1-4B The Reflex-1 skills (LoRA + latent modules) and registry. Needed for every variant 1.8 GB —
matu79go/Reflex-1-4B-bnb-4bit (this page) Base model, pre-quantized 4-bit NF4. Recommended 9.3 GB ~10 GB
google/gemma-4-E4B-it Base model, original BF16 (quantize at load with --quant nf4, or run in BF16 with --quant none) ~16 GB ~10 GB / ~16 GB

Open In Colab

Usage

git clone https://github.com/matu79go/reflex-1.git && cd reflex-1
pip install -r requirements.txt
python -m reflex.server --base matu79go/Reflex-1-4B-bnb-4bit --skills skills.json --port 8097

Results with this base are the same as quantizing google/gemma-4-E4B-it at load time (checked on intent classification, image classification and the three reasoning skills).

License

Apache License 2.0. Derived from Google Gemma 4 E4B (Apache 2.0); only quantized, no other changes. Not affiliated with or endorsed by Google.

Downloads last month
24
Safetensors
Model size
8B params
Tensor type
BF16
·
U8
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for matu79go/Reflex-1-4B-bnb-4bit

Quantized
(368)
this model

Collection including matu79go/Reflex-1-4B-bnb-4bit