Gemma4-12B-it-Fast

A calibrated mixed NVFP4/FP8 export of DevelopingDad's Gemma 4 12B IT Trial62, with an optional vLLM serving recipe tested on one NVIDIA DGX Spark GB10. This release contains the same target weights used in the October 8, 2026 API experiments.

Primary credit: Google DeepMind / the Gemma Team created the original Gemma model family, architecture, instruction-tuned model, tokenizer, processor and native chat format. DevelopingDad supplied the immediate Trial62 checkpoint, produced this quantized derivative, and conducted the serving experiments. The separately downloaded Google assistant model supplies speculative drafts in the measured fast recipe.

Release status: experimental. The final short, single-stream API check measured 52.98 tok/s median / 50.55 minimum. These rates use the full optional MTP8 recipe and the response-contract template described below. Full quality qualification remains unmet, and a uniform 40 tok/s floor across long contexts and concurrent requests was not established.

Provenance and changes

Item Identity
Immediate source DevelopingDad/gemma-4-12b-it-trial62
Source revision 568aa9764b4f695fb18881b97f6d4ee8b725f75c
Upstream family / original IT model Google DeepMind, google/gemma-4-12B-it
Derivative Calibrated mixed NVFP4 MLPs / FP8 attention, export V5
Target weight file model.safetensors, 9,304,966,064 bytes
Original target embedding/head Tied BF16, retained byte-for-byte
Native tokenizer / processor / chat template Preserved from Trial62
Additional optimization training None in this conversion

The source contains 11,959,730,224 stored BF16 parameter elements across 48 language layers. Its native configuration permits 262,144 total positions; this release's tested serving recipe uses 32,768 total prompt-plus-output tokens.

The Hub's automatic parameter counter can count packed U8 storage elements rather than logical weights. Each NVFP4 byte stores two logical weights; the smaller automatic count does not mean this is an 8B architecture.

Text MLP weights use importance-weighted MSE NVFP4, groups of 16, static FP8 block scales and global divisors. MLP activations use locally dynamic NVFP4. Attention weights use per-channel FP8 and activations use dynamic per-token FP8. K/V scales are calibrated FP8 scalars. Original tied BF16 embedding/output-head and media projections are retained.

Calibration used 128 English C4 training rows, seed 91027, maximum 1,024 tokens, a 512-row shuffle buffer, and dataset revision 607bd4c8450a42878aa9ddc051a65a055450ef87. Benchmark/quality prompts and observed answers were excluded from calibration. This was post-training quantization; no new alignment or refusal-removal training was performed. The original Trial62 modification method is not documented in its source repository and is not inferred here.

The converter handles Gemma's heterogeneous attention dimensions per layer: 40 sliding layers with eight KV heads and eight full layers with one KV head. Gate/up projection observers share their global statistics before packing. All 48 KV initializers, 144 packed MLP weights, and 48 shared gate/up scale pairs passed the export audits. These checks establish export integrity, not broad quality equivalence.

See MODIFICATIONS.md, provenance.json, and SHA256SUMS.

Credits and attribution

Contribution Credit and source
Original model, architecture, IT training, tokenizer, processor, native template Google DeepMind / Gemma Team, Gemma 4 12B IT, technical report
Immediate Trial62 checkpoint and this release's conversion, benchmarking, response contract and integration DevelopingDad, Trial62
Native speculative assistant Google DeepMind / Gemma Team, Gemma 4 12B IT assistant, pinned revision 46d4c6f13f0ac0ad827b915669b8df9b81c64c51
Serving engine and upstream code extended by the optional overlay vLLM contributors, vLLM, source db9527a46873454610df6dbedf79a36d6bf1a7f6
Quantization and compressed checkpoint tooling vLLM Project / Red Hat AI and the llm-compressor and compressed-tensors contributors, llm-compressor, compressed-tensors
Calibration corpus Google/T5 and the Allen Institute for AI, C4, derived from Common Crawl; C4 dataset license is ODC-BY
GPU kernels, framework and hardware FlashInfer contributors, FlashInfer; NVIDIA, CUTLASS, CUDA and DGX Spark; PyTorch contributors, PyTorch; Hugging Face contributors, Transformers
Preceding reference checkpoints used to compare serving recipes Unsloth, Gemma 4 12B IT NVFP4, and Red Hat AI, Gemma 4 12B IT NVFP4

DevelopingDad used OpenAI Codex assistance for implementation, audits and documentation. DevelopingDad maintains and publishes this release.

The immediate source for this export is Trial62. The reference publishers' weights were not merged into this checkpoint. Google assistant weights are downloaded separately and are not bundled here. No upstream organization is represented as endorsing this derivative.

Tested vLLM recipe

The benchmark environment was ARM64, one DGX Spark GB10, vLLM 0.31.0, Torch 2.13.0+cu130, Transformers 5.17.0, FlashInfer 0.7.0.post1, and compressed-tensors 0.17.0. Conversion separately used llm-compressor 0.14.0 / compressed-tensors 0.19.0 / datasets 5.0.1.

  • Public container reference: vllm/vllm-openai@sha256:3f7dd5b777d34d1724456ce71f87385dca288c3bb23029ab27dee358f5d2b971.
  • Measured ARM64 local image ID: sha256:bbe7045055707d1027079ed01cca40812578a235c56032d85c16108ad473384c; its binding to the public container reference is recorded in provenance.json.
  • Eight speculative tokens, Google's pinned assistant, FP8-per-tensor draft body, and the experimental isolated FP8 draft-head overlay.
  • BF16 unquantized target components; 8 GiB FP8 KV, four sequence slots, 4,096 batched tokens, chunked prefill, prefix caching off.
  • Triton attention for the heterogeneous 256/512 head dimensions; FlashInfer CUTLASS NVFP4 linear kernel.
  • Thinking off by default; native Gemma 4 reasoning and tool parsers retained.
  • Normal sampling: temperature 1.0, top-p 0.95, top-k 64, seed 91027, native EOS, 2,048-token benchmark output ceiling.

The optional overlay quantizes only a separate draft output head; the target's tied BF16 embedding/head remains protected. A post-share startup receipt checks dtype, scale, shape and aliasing. A per-boot archival wrapper preserves older receipts before each start. These are experimental extensions of the pinned runtime.

The root chat_template.jinja is the preserved native template. The measured recipe explicitly selects serving/gemma4-profile-chat-template.jinja, which prepends a general source-fidelity response contract. Its extra prompt tokens and output-policy effects are part of these measurements.

Download the repository, then follow serving/README.md. The included launcher prints its plan by default and requires --run to start a new loopback-bound server. Using the target checkpoint with another engine, template, GPU, speculation setting or context length requires separate validation.

Example request to a locally started server:

from openai import OpenAI

client = OpenAI(base_url="http://127.0.0.1:8000/v1", api_key="unused")
response = client.chat.completions.create(
    model="DevelopingDad/Gemma4-12B-it-Fast",
    messages=[{"role": "user", "content": "Explain how a library can keep lending books during a catalog migration."}],
    temperature=1.0,
    top_p=0.95,
    max_tokens=1024,
    extra_body={"top_k": 64, "chat_template_kwargs": {"enable_thinking": False}},
)
print(response.choices[0].message.content)

Set chat_template_kwargs.enable_thinking=true to request reasoning. Use explicit JSON schemas for strict structured output; plain JSON-only prompting has known fence failures.

Measured performance

Five frozen mixed prose/code prompts were tested at one stream (c1) and four concurrent streams (c4). The final local-client HTTP check through HTTPS completed five measured short c1 requests: 52.98 / 50.55 tok/s median/minimum conservative visible delivery, 0.187 s median first text, 55.42 generated tok/s aggregate, all five native EOS, no caps or errors. One 16-token diagnostic warmup was excluded.

Warm normal workload Rows Per-stream delivery median / minimum tok/s Median first text seconds Aggregate generated tok/s Native EOS / cap
c1 short, median 857 prompt tokens 5 52.68 / 50.15 0.205 54.93 5 / 0
c1 ~8K, median 8,216 prompt tokens 5 54.04 / 52.57 2.127 47.68 5 / 0
c1 near 30K, median 30,390 prompt tokens 5 46.03 / 38.93 12.622 22.91 5 / 0
c4 short 20 49.12 / 35.72 0.608 189.22 20 / 0
c4 ~8K 20 44.10 / 22.84 6.064 104.68 20 / 0
c4 near 30K 20 18.28 / 3.82 33.092 35.39 19 / 1

The 75-row normal run retained 74 native stops, one 2,048-token cap, and zero transport errors. The capped request was a c4 near-30K migration-plan response. Conservative delivery uses the interval from first visible content to SSE completion, subtracting the first fragment and a terminal allowance. Aggregate generation uses endpoint usage, including EOS/control-token accounting, over homogeneous request-wave intervals and includes prefill; it is not per-stream decode speed. Concurrent output can pause while other requests prefill.

A separate temperature-0 short screen measured c1 median/minimum 58.08/42.02 and c4 53.59/31.89, with 211.60 aggregate generated tok/s and 25 native stops. Sampling and model differences prevent treating earlier reference-publisher timings as a runtime-only improvement. This release does not claim global fastest-possible optimality.

Benchmark definitions and summaries document the workload and calculation boundaries.

Quality, capabilities and known limits

Two original direct-model content runs at c1/c4 each had 46 mechanical passes and two application-dependent cases unscored, with all 48 responses stopping naturally. Raw semantic review found 40 clean, one minor issue, and five hard JSON-fence failures among the 46 scoreable cases. The mechanical JSON grader strips fences, so its pass count alone is insufficient.

A six-case supplemental holdout had four mechanical passes and two failures. The derivative corrected the source's invoice arithmetic, but retained an untrusted-text failure that repeated a prohibited phrase. A semantically correct transfer answer also missed a narrow frozen wording matcher. These failures remain recorded; the grader was not relaxed.

Native strict JSON-schema output, tool call, tool-result continuation and opt-in reasoning passed direct API checks. Eight fixed benign willingness probes were willing, usable and naturally stopped for source and derivative. This is bounded behavior evidence, not proof that quantization preserves every behavior or that refusals are universally removed.

Media-related components and native protocols are retained, but this release did not complete a new end-to-end image/audio/video qualification. The original application's retrieval, authorization, citations, persistence and client behavior were not exercised. Full qualification is incomplete. The currently published artifact precedes further issue-fix experiments.

License, notices and citation

The model derives from Google's Apache-2.0 Gemma 4 release. This repository includes LICENSE, NOTICE, and MODIFICATIONS.md. Existing upstream SPDX/copyright headers are retained in the optional runtime source files. C4's separate ODC-BY license and Common Crawl terms apply to the calibration corpus; no corpus text is distributed here.

Please credit the original Gemma authors, DevelopingDad's Trial62 checkpoint, this conversion, and any assistant/runtime/tooling used in your deployment. Google's original citation is:

@misc{gemmateam2026gemma4,
  title={Gemma 4 Technical Report},
  author={Gemma Team},
  year={2026},
  eprint={2607.02770},
  archivePrefix={arXiv},
  primaryClass={cs.CL},
  url={https://arxiv.org/abs/2607.02770}
}
Downloads last month
9
Safetensors
Model size
8B params
Tensor type
BF16
·
F8_E4M3
·
U8
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for DevelopingDad/Gemma4-12B-it-Fast

Quantized
(1)
this model

Dataset used to train DevelopingDad/Gemma4-12B-it-Fast

Paper for DevelopingDad/Gemma4-12B-it-Fast