Instructions to use unsloth/MiniMax-H3-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Local Apps Settings
- Unsloth Studio
How to use unsloth/MiniMax-H3-GGUF with Unsloth Studio:
Install Unsloth Studio (macOS, Linux, WSL)
curl -fsSL https://unsloth.ai/install.sh | sh # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for unsloth/MiniMax-H3-GGUF to start chatting
Install Unsloth Studio (Windows)
irm https://unsloth.ai/install.ps1 | iex # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for unsloth/MiniMax-H3-GGUF to start chatting
Using HuggingFace Spaces for Unsloth
# No setup required # Open https://huggingface.co/spaces/unsloth/studio in your browser # Search for unsloth/MiniMax-H3-GGUF to start chatting
Load model with FastModel
pip install unsloth from unsloth import FastModel model, tokenizer = FastModel.from_pretrained( model_name="unsloth/MiniMax-H3-GGUF", max_seq_length=2048, )
Is the arch labeled as minimax?
I noticed people commenting on my posts about these ggufs and they said the arch doesnt pass through the gguf loader and is causing errors for them.
When i qusnted the model i passed it as "wan" arch and keys to avoid this error and alot of the time you can trick the quantizer to believe its wan then swap architecture back later but i left it at wan because it passes checks in the gguf loader without needing to code for it and PR it on github!
I do have a gguf loader as well that supports the model passing as wan for if you want to do so:
(It says w3a8 loader but thats just the main use, it does gguf and w4a8 if need be but thats native now)
https://github.com/RealRebelAI/Rebels_w3a8_Loader