Gemma 4 31B Spliced GGUF

A highly optimized, spliced variant of Google's flagship Gemma 4 31B-it model.

Model Description

This repository contains a spliced mixed-precision GGUF binary of the Gemma 4 31B-it architecture. By removing redundant layers, the active parameter size is reduced, enabling efficient local inference on unified memory architectures (like standard consumer Apple Silicon laptops).

  • Format: GGUF (Q4_K_M / IQ4_XS)
  • Sizes: 17.84 GB (Q4_K_M) / 16.10 GB (IQ4_XS)
  • Target Platforms: Apple Silicon MacBooks (M1/M2/M3/M4) and standard CPU/GPU local runtimes.

Local Quickstart

Run the model natively via standard llama-cli:

# For standard execution:
llama-cli -ngl 100 -m ./gemma4_spliced_q4km.gguf -p "The mathematical beauty of wavelets lies in" -n 128

# For high-efficiency, low-VRAM execution (Recommended):
llama-cli -ngl 100 -m ./gemma4_spliced_iq4xs.gguf -p "The mathematical beauty of wavelets lies in" -n 128
Downloads last month
59
GGUF
Model size
29B params
Architecture
gemma4
Hardware compatibility
Log In to add your hardware

We're not able to determine the quantization variants.

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support