Instructions to use huihui-ai/Huihui-gemma-4-12B-it-qat-q4_0-unquantized-abliterated with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use huihui-ai/Huihui-gemma-4-12B-it-qat-q4_0-unquantized-abliterated with Transformers:
# Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("huihui-ai/Huihui-gemma-4-12B-it-qat-q4_0-unquantized-abliterated") model = AutoModelForMultimodalLM.from_pretrained("huihui-ai/Huihui-gemma-4-12B-it-qat-q4_0-unquantized-abliterated") - Notebooks
- Google Colab
- Kaggle
would there be one for 25b-a4b ?
or other QAT gemma from google ?
There is QAT for every version, I can only assume they are working on uncensoring the others, really excited to get my hands on 31B version
Yes, please wait patiently.
Also can u make Stepfun flash 3.7 pls?
Also can u make Stepfun flash 3.7 pls?
Will give it a try later.
thanks for the 26b !
I'm able to do 49K context on my 16 GB card with the official Gemma QAT 31B weights on Ollama, but with the uncensored QAT weights I had to lower it or get out of memory errors weirdly, might be Ollama's fault partially for not calculating memory usage very well and just offloading correctly but it seems the model is slightly larger than Google's QAT version.
Also, my main use-case is creative writing, and maybe just me, but it feels a bit different than normal QAT Gemma much more of slop feeling. Overall though, model is nice, and at least feels smarter than normal old Q4