JSQ โ€” Qwen3-8B-Base (50% unstructured sparsity + INT4, symmetric)

Baseline compressed checkpoint. Method: JSQ. Sparsity: 50% unstructured. Quantization: INT4, per-group 128, symmetric. Base: Qwen/Qwen3-8B-Base.

  • model*.safetensors: sparse + fake-quant weights (fp16).
  • compression/scales.safetensors: per-group-128 symmetric scales [out, in/128].
  • compression_config.json: method / sparsity / bits / granularity / symmetric. Mask = weight == 0; symmetric levels in [-7,7].
Downloads last month
-
Safetensors
Model size
8B params
Tensor type
BF16
ยท
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support