Model Details

This model is an int4 model with group_size 128 and symmetric quantization of Qwen/Qwen3.5-4B generated by [intel/auto-round]. Please follow the license of the original model.

Cite

@article{cheng2023optimize,
  title={Optimize weight rounding via signed gradient descent for the quantization of llms},
  author={Cheng, Wenhua and Zhang, Weiwei and Shen, Haihao and Cai, Yiyang and He, Xin and Lv, Kaokao and Liu, Yi},
  journal={arXiv preprint arXiv:2309.05516},
  year={2023}
}

arxiv github

Downloads last month
4
Safetensors
Model size
2B params
Tensor type
I32
·
BF16
·
F16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for bachngo/Qwen3.5-4B-int4-AR

Finetuned
Qwen/Qwen3.5-4B
Quantized
(401)
this model

Paper for bachngo/Qwen3.5-4B-int4-AR