gemma-2-2b โ€” FP8_DYNAMIC (W8A8-e4m3)

Weight-FP8 checkpoint of google/gemma-2-2b-it, produced for the TR171 deployment-time safety-tax benchmark.

Provenance

Field Value
Base model google/gemma-2-2b-it
Base revision 299a8560bedf22ed1c72a8a11e7dce4a7f9f51f8 (verified โ€” recorded in the frozen matrix)
Recipe FP8_DYNAMIC (W8A8-e4m3), llmcompressor
Quantization method compressed-tensors
Calibration data none โ€” FP8_DYNAMIC is data-free
Build date 2026-07-02
Shard size 3.24 GB
Quantize wall time 177.5 s

Reproducing

Producer: research/tr171/expansion/fp8_support_probe.py; environment: research/tr171/expansion/Dockerfile.fp8. The recipe takes no calibration corpus, so there is no dataset or seed to reproduce โ€” only the base checkpoint and the toolchain version.

Known reproducibility gap: llmcompressor was unpinned at build time, so the exact version used on 2026-07-02 is unrecorded. The Dockerfile now pins it. A rebuild may therefore not be bit-identical to this artifact; the sha256 recorded in fp8_support_matrix.json will detect a difference but cannot repair one.

License and notices

Gemma is provided under and subject to the Gemma Terms of Use found at ai.google.dev/gemma/terms. Use is subject to the Gemma Prohibited Use Policy at ai.google.dev/gemma/prohibited_use_policy. These terms travel with this derivative and must be passed on to any downstream recipient.

This FP8 derivative inherits the upstream terms of google/gemma-2-2b-it. Consult the base model's licence before redistributing.

Downloads last month
-
Safetensors
Model size
3B params
Tensor type
BF16
ยท
F8_E4M3
ยท
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for Crusadersk/gemma-2-2b-it-FP8-Dynamic-TR171

Quantized
(192)
this model

Collection including Crusadersk/gemma-2-2b-it-FP8-Dynamic-TR171