Instructions to use OptimizeLLM/Qwen3.8-27B-heretic-MTP-BF16 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use OptimizeLLM/Qwen3.8-27B-heretic-MTP-BF16 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="OptimizeLLM/Qwen3.8-27B-heretic-MTP-BF16") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("OptimizeLLM/Qwen3.8-27B-heretic-MTP-BF16") model = AutoModelForMultimodalLM.from_pretrained("OptimizeLLM/Qwen3.8-27B-heretic-MTP-BF16", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use OptimizeLLM/Qwen3.8-27B-heretic-MTP-BF16 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "OptimizeLLM/Qwen3.8-27B-heretic-MTP-BF16" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "OptimizeLLM/Qwen3.8-27B-heretic-MTP-BF16", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/OptimizeLLM/Qwen3.8-27B-heretic-MTP-BF16
- SGLang
How to use OptimizeLLM/Qwen3.8-27B-heretic-MTP-BF16 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "OptimizeLLM/Qwen3.8-27B-heretic-MTP-BF16" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "OptimizeLLM/Qwen3.8-27B-heretic-MTP-BF16", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "OptimizeLLM/Qwen3.8-27B-heretic-MTP-BF16" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "OptimizeLLM/Qwen3.8-27B-heretic-MTP-BF16", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Docker Model Runner
How to use OptimizeLLM/Qwen3.8-27B-heretic-MTP-BF16 with Docker Model Runner:
docker model run hf.co/OptimizeLLM/Qwen3.8-27B-heretic-MTP-BF16
Qwen3.8-27B-heretic-MTP-BF16
An abliterated build of Qwen/Qwen3.8-27B in BF16, with the MTP heads and full vision tower intact. This is the source checkpoint. Quantize it to whatever format you need.
For the FP8 build, the measured numbers, the flags, and the full method writeup, see OptimizeLLM/Qwen3.8-27B-heretic-MTP-FP8.
- Base: Qwen3.8-27B (27.8B dense, 64 layers, hybrid linear/full attention, native image and video).
- Abliteration: HERETIC 1.4.0 with local patches. Two-stage slot-grouped pipeline. KL 0.065 against base on harmless prompts.
- Refusals: one out of 99 for the base model, on a 100-prompt harmful-behaviors set reviewed manually.
- Non-refusing, not neutral. Directional ablation removes the tendency to decline, not a trained-in viewpoint. It will engage with any topic; it can still reflect the base model's lean on one.
- Tokenizer: official Qwen3.8-27B's, byte-identical to upstream.
The reproduce/ folder holds the HERETIC config, seed parameters, and the local patches.
You are responsible for what you do with it.
Thanks
- Qwen Team, for Qwen3.8-27B and the open release.
- p-e-w, for HERETIC.
- The community that first worked out the two-stage slot-grouped MPOA recipe on the Qwen3.6 family.
License
Inherits the base Qwen3.8-27B license. Apache-2.0.
- Downloads last month
- 84
Model tree for OptimizeLLM/Qwen3.8-27B-heretic-MTP-BF16
Base model
Qwen/Qwen3.8-27B