Instructions to use Minachist/Muse-Glimmer-30B-INT8-AutoRound with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Minachist/Muse-Glimmer-30B-INT8-AutoRound with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="Minachist/Muse-Glimmer-30B-INT8-AutoRound") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("Minachist/Muse-Glimmer-30B-INT8-AutoRound") model = AutoModelForMultimodalLM.from_pretrained("Minachist/Muse-Glimmer-30B-INT8-AutoRound", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use Minachist/Muse-Glimmer-30B-INT8-AutoRound with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Minachist/Muse-Glimmer-30B-INT8-AutoRound" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Minachist/Muse-Glimmer-30B-INT8-AutoRound", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/Minachist/Muse-Glimmer-30B-INT8-AutoRound
- SGLang
How to use Minachist/Muse-Glimmer-30B-INT8-AutoRound with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "Minachist/Muse-Glimmer-30B-INT8-AutoRound" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Minachist/Muse-Glimmer-30B-INT8-AutoRound", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "Minachist/Muse-Glimmer-30B-INT8-AutoRound" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Minachist/Muse-Glimmer-30B-INT8-AutoRound", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Docker Model Runner
How to use Minachist/Muse-Glimmer-30B-INT8-AutoRound with Docker Model Runner:
docker model run hf.co/Minachist/Muse-Glimmer-30B-INT8-AutoRound
another models quantization
can you please quantize to INT8 symmetric two models: bottlecapai/ThinkingCap-Qwen3.6-27B and endless-frontier/BigBang-v1.
I would really like to test them all. my hardware is 2x3090 with nvlink. i am already use your Qwen3.6-27B-INT8-Autoround-V2 as for everyday tasks and it is very good.
Minachist/Qwen3.6-27B-INT8-Autoround-V2 already ships the quantization script, and usage is documented in the repo, so please start there. If anything in it is unclear, feel free to ask an LLM.
ThinkingCap-Qwen3.6-27B is just a finetune of Qwen3.6, so that script should work on it directly.For BigBang-v1 (35B base), please don't use the older 35B script. The V2 script is more refined, so it's better to adapt that one instead. The 35B script exists mainly for cases where editing V2 isn't worth the effort.I don't generally do quantization for finetuned models as they are not really good. Ornith was an exception as it seemed promising (and I wanted to add MTP layers for fun). I'd like you to run these yourself using the above. If you hit any issues, feel free to ask here.
Update: I changed my mind. Once Qwen3.8-27B is out I'll have some free time, so I can quantize and publish BigBang-v1 quantization myself. ThinkingCap-Qwen3.6-27B is gated (despite the Apache-2.0 license) and since I'm not handing over personal info to an unknown party to request access, so that one's off the table permanently.