oQ4e quantization request

#1
by Izenflamm - opened

Hi! Please, could you quantize the model to oQ4e and oQ5e through oMLX? Not many people have the hardware needed to do it, but 128GB unified memory Mac users could use an optimized quantization of the model for local inference. Thank you!

Vontra org

Leave it with me.

Vontra org
ashxhart changed discussion status to closed

Sign up or log in to comment