llm-quant-artifacts

  • What: fake-quantized LLM checkpoints (weights 4-bit, group 128) used in a quantization study, packed losslessly.
  • Use: transfer between our own machines; restore with the companion tool, which decodes .lqz back to the original model.pt and checks every tensor's sha256 against manifest.json.
  • Layout
    • <model>/g_<method>/model.lqz: one packed checkpoint per quantization method
    • <model>/g_qoq/smooth.pt: per-channel smoothing scales used by that method at runtime
    • <model>/bases/, <model>/pca/: per-layer key-side bases and weight-side PCA bases used at runtime
    • <model>/manifest.json: file list, sha256 of every file and tensor, restore paths
  • Format (.lqz): a 4-bit group of 128 values holds at most 16 distinct fp16 values, stored as a 16-entry fp16 table plus 128 4-bit indices; all other tensors are stored raw. Decoding is exact (bit-identical tensors).
  • Not an inference-ready release: research artifacts only, no quality guarantee.

Models and licenses

Each folder carries the license of the model it is derived from.

Folder Derived from License
q3_4b/ Qwen/Qwen3-4B Apache-2.0
l31i/ meta-llama/Llama-3.1-8B-Instruct Llama 3.1 Community License (l31i/LICENSE, l31i/USE_POLICY.md)
l2/ meta-llama/Llama-2-7b-hf Llama 2 Community License (l2/LICENSE.txt, l2/USE_POLICY.md)

Built with Llama. Llama 3.1 is licensed under the Llama 3.1 Community License, Copyright © Meta Platforms, Inc. All Rights Reserved. Llama 2 is licensed under the LLAMA 2 Community License, Copyright (c) Meta Platforms, Inc. All Rights Reserved. Use of the l31i/ and l2/ artifacts is subject to those licenses and to Meta's Acceptable Use Policies included alongside.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support