Model Details

This model is an int4 model with group_size 128 and symmetric quantization of Qwen/Qwen3.5-0.8B generated by [intel/auto-round]. Please follow the license of the original model.

Cite

@article{cheng2023optimize,
  title={Optimize weight rounding via signed gradient descent for the quantization of llms},
  author={Cheng, Wenhua and Zhang, Weiwei and Shen, Haihao and Cai, Yiyang and He, Xin and Lv, Kaokao and Liu, Yi},
  journal={arXiv preprint arXiv:2309.05516},
  year={2023}
}

arxiv github

Downloads last month
3
Safetensors
Model size
0.5B params
Tensor type
I32
·
BF16
·
F16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for bachngo/Qwen3.5-0B-int4-AR

Quantized
(234)
this model

Paper for bachngo/Qwen3.5-0B-int4-AR