Suggesting re-quantization of this model from source to OptiQ-4bit

#2
by ahmedihamdy - opened

Hello Folks,

As always thanks for all the Great work and the helpful model files,

Anecdotally, I feel this version could benefit from reprocessing or re-quantization into the exact same 4bit Optiq, perhaps with a quality first setting.

Anecdotally, I observe it is more prone to skip instructions or to go into endless loops while thinking as compared to other four-bit models that are also based on the Quen architecture, the 6 bit OptiQ version however Runs perfectly with immense observed difference in instruction following and general quality of outputs.

I know this is an anecdotal observation but wanted to share it,if you have time to reprocess this file with perhaps the latest Python versions from OpitQ and the latest quality focused settings.

Thank you.

MLX Community org

Thanks for the detailed feedback and for sharing your observations. The comparison with the 6-bit OptiQ version is particularly useful. We’ll take a closer look at the quantization process and instruction-following behavior, and consider reprocessing it with the latest OptiQ tooling and quality-focused settings if we can reproduce the issue.

MLX Community org

We'll investigate further

Thank you Siddh and jorgemunozl, many people will benefit from running models locally who can't afford to pay the monthly premium.

MLX Community org

That’s the mission — putting capable AI back in the hands of the people, locally and without the paywall.

Sign up or log in to comment