Since nobody appears to want to post Q8_0 mmproj and the default is over 5gb in size, it's hard to get all your context into 96gb. Downloading the weights is a nonstarter because they are 200G+
All facilities for requantizing mmproj are broken in both ik_llama.cpp and llama.cpp. What do?
Screw around with models and use python! Took a while but now q8_0 mmproj exists and works. Maybe also the script is useful for someone?
- Downloads last month
- 18
Hardware compatibility
Log In to add your hardware
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐ Ask for provider support
Model tree for Lockout/Mistral-Medium-3.5-128B-GGUF-Q8_0-mmproj
Base model
mistralai/Mistral-Medium-3.5-128B