DiffGemma 26B-A4B-it β€” q4 (.dgq)

A quantized, self-contained .dgq pack of DiffusionGemma 26B-A4B-it (Gemma-4 26B-A4B MoE, discrete block diffusion) for the diffgemma-mps Rust + Metal inference engine on Apple Silicon. Text-only (v1).

Quantization β€” q4 profile

Tensor class Format Bits/weight
MoE experts q4_block (int4 affine, group 32) 5.0
Attention (q/k/v/o), dense FFN bf16 16
Embeddings, router, norms bf16 16
Self-conditioning MLP q8_row 8

~19 GiB of weights. Manifest version 1 (affine) β€” loads on every diffgemma-mps build.

Usage

diffgemma-mps download --repo mmastrac/diffgemma-26b-a4b-it-q4 -o model/diffgemma-26b-a4b-it-q4
diffgemma-mps -m model/diffgemma-26b-a4b-it-q4 chat

Build and run details: github.com/mmastrac/diffgemma.

License

Gemma Terms of Use. Derived from google/diffusiongemma-26B-A4B-it; use is subject to the Gemma license.

Downloads last month
43
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for mmastrac/diffgemma-26b-a4b-it-q4

Finetuned
(20)
this model