llm-quant-artifacts
- What: fake-quantized LLM checkpoints (weights 4-bit, group 128) used in a quantization study, packed losslessly.
- Use: transfer between our own machines; restore with the companion tool, which decodes
.lqzback to the originalmodel.ptand checks every tensor's sha256 againstmanifest.json. - Layout
<model>/g_<method>/model.lqz: one packed checkpoint per quantization method<model>/g_qoq/smooth.pt: per-channel smoothing scales used by that method at runtime<model>/bases/,<model>/pca/: per-layer key-side bases and weight-side PCA bases used at runtime<model>/manifest.json: file list, sha256 of every file and tensor, restore paths
- Format (.lqz): a 4-bit group of 128 values holds at most 16 distinct fp16 values, stored as a 16-entry fp16 table plus 128 4-bit indices; all other tensors are stored raw. Decoding is exact (bit-identical tensors).
- Not an inference-ready release: research artifacts only, no quality guarantee.
Models and licenses
Each folder carries the license of the model it is derived from.
| Folder | Derived from | License |
|---|---|---|
q3_4b/ |
Qwen/Qwen3-4B |
Apache-2.0 |
l31i/ |
meta-llama/Llama-3.1-8B-Instruct |
Llama 3.1 Community License (l31i/LICENSE, l31i/USE_POLICY.md) |
l2/ |
meta-llama/Llama-2-7b-hf |
Llama 2 Community License (l2/LICENSE.txt, l2/USE_POLICY.md) |
Built with Llama. Llama 3.1 is licensed under the Llama 3.1 Community License, Copyright © Meta Platforms, Inc. All Rights
Reserved. Llama 2 is licensed under the LLAMA 2 Community License, Copyright (c) Meta Platforms, Inc. All Rights Reserved.
Use of the l31i/ and l2/ artifacts is subject to those licenses and to Meta's Acceptable Use Policies included alongside.
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support