DiffGemma 26B-A4B-it β€” nvfp4 (.dgq)

A quantized, self-contained .dgq pack of DiffusionGemma 26B-A4B-it (Gemma-4 26B-A4B MoE, discrete block diffusion) for the diffgemma-mps Rust + Metal inference engine on Apple Silicon. Text-only (v1).

The most compressed profile: MoE experts and attention + dense FFN are all NVFP4.

Quantization β€” nvfp4 profile

Tensor class Format Bits/weight
MoE experts nvfp4_block (NVFP4, group 16) ~4.5
Attention (q/k/v/o), dense FFN nvfp4_block (NVFP4, group 16) ~4.5
Embeddings, router, norms bf16 16
Self-conditioning MLP q8_row 8

~15 GiB of weights. Manifest version 2 (nvfp4) β€” loads on any diffgemma-mps build with NVFP4 support.

Usage

diffgemma-mps download --repo mmastrac/diffgemma-26b-a4b-it-nvfp4 -o model/diffgemma-26b-a4b-it-nvfp4
diffgemma-mps -m model/diffgemma-26b-a4b-it-nvfp4 chat

Build and run details: github.com/mmastrac/diffgemma.

License

Gemma Terms of Use. Derived from google/diffusiongemma-26B-A4B-it; use is subject to the Gemma license.

Downloads last month
14
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for mmastrac/diffgemma-26b-a4b-it-nvfp4

Finetuned
(20)
this model