Built with Llama

llama3.2-3b — FP8_DYNAMIC (W8A8-e4m3)

Weight-FP8 checkpoint of unsloth/Llama-3.2-3B-Instruct, produced for the TR171 deployment-time safety-tax benchmark.

Provenance

Field Value
Base model unsloth/Llama-3.2-3B-Instruct
Base revision 006f5dcd1393c3add266de40994ba96225e9689d (INFERRED — recovered from local HF cache snapshot, not a recorded fact)
Recipe FP8_DYNAMIC (W8A8-e4m3), llmcompressor
Quantization method compressed-tensors
Calibration data none — FP8_DYNAMIC is data-free
Build date 2026-07-02
Shard size 3.63 GB
Quantize wall time 37.7 s

Reproducing

Producer: research/tr171/expansion/fp8_support_probe.py; environment: research/tr171/expansion/Dockerfile.fp8. The recipe takes no calibration corpus, so there is no dataset or seed to reproduce — only the base checkpoint and the toolchain version.

Known reproducibility gap: llmcompressor was unpinned at build time, so the exact version used on 2026-07-02 is unrecorded. The Dockerfile now pins it. A rebuild may therefore not be bit-identical to this artifact; the sha256 recorded in fp8_support_matrix.json will detect a difference but cannot repair one.

License and notices

Llama 3.2 is licensed under the Llama 3.2 Community License, Copyright (c) Meta Platforms, Inc. All Rights Reserved.

This FP8 derivative inherits the upstream terms of unsloth/Llama-3.2-3B-Instruct. Consult the base model's licence before redistributing.

Downloads last month
-
Safetensors
Model size
3B params
Tensor type
BF16
·
F8_E4M3
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Crusadersk/Llama-3.2-3B-Instruct-FP8-Dynamic-TR171

Quantized
(113)
this model

Collection including Crusadersk/Llama-3.2-3B-Instruct-FP8-Dynamic-TR171