YAML Metadata Warning:empty or missing yaml metadata in repo card

Check out the documentation for more information.

Optimized for resource-constrained and mobile deployment through targeted architectural and numerical optimizations.

Audio encoders removed to reduce the overall model footprint.

Vision layers quantized to Q8_0 or Q4_K_M for substantially reduced storage requirements.

Positional embeddings converted from FP32 to FP16 to further reduce model size while retaining higher precision than the quantized vision layers.

These optimizations reduce the mmproj-BF16 file size to approximately 14%/20% of the original for the Q4_K_M / Q8_0 respectively, significantly lowering the storage requirements for on-device deployment.

license: gemma

Downloads last month
44
GGUF
Model size
5B params
Architecture
gemma4
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support