Qwen3.8 MLX Quants
Collection
MLX affine 4-bit, 6-bit, and 8-bit quantizations of Qwen/Qwen3.8-27B. • 3 items • Updated
How to use DreamFoundries/Qwen3.8-27B-8bit with MLX:
# Make sure mlx-lm is installed
# pip install --upgrade mlx-lm
# Generate text with mlx-lm
from mlx_lm import load, generate
model, tokenizer = load("DreamFoundries/Qwen3.8-27B-8bit")
prompt = "Write a story about Einstein"
messages = [{"role": "user", "content": prompt}]
prompt = tokenizer.apply_chat_template(
messages, add_generation_prompt=True
)
text = generate(model, tokenizer, prompt=prompt, verbose=True)How to use DreamFoundries/Qwen3.8-27B-8bit with Pi:
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "DreamFoundries/Qwen3.8-27B-8bit"
# Install Pi:
npm install -g @mariozechner/pi-coding-agent
# Add to ~/.pi/agent/models.json:
{
"providers": {
"mlx-lm": {
"baseUrl": "http://localhost:8080/v1",
"api": "openai-completions",
"apiKey": "none",
"models": [
{
"id": "DreamFoundries/Qwen3.8-27B-8bit"
}
]
}
}
}# Start Pi in your project directory: pi
How to use DreamFoundries/Qwen3.8-27B-8bit with OpenClaw:
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "DreamFoundries/Qwen3.8-27B-8bit"
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "DreamFoundries/Qwen3.8-27B-8bit" \ --custom-provider-id mlx-lm \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
openclaw agent --local --agent main --message "Hello from Hugging Face"
How to use DreamFoundries/Qwen3.8-27B-8bit with MLX LM:
# Install MLX LM uv tool install mlx-lm # Interactive chat REPL mlx_lm.chat --model "DreamFoundries/Qwen3.8-27B-8bit"
# Install MLX LM
uv tool install mlx-lm
# Start the server
mlx_lm.server --model "DreamFoundries/Qwen3.8-27B-8bit"
# Calling the OpenAI-compatible server with curl
curl -X POST "http://localhost:8000/v1/chat/completions" \
-H "Content-Type: application/json" \
--data '{
"model": "DreamFoundries/Qwen3.8-27B-8bit",
"messages": [
{"role": "user", "content": "Hello"}
]
}'How to use DreamFoundries/Qwen3.8-27B-8bit with Hermes Agent:
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "DreamFoundries/Qwen3.8-27B-8bit"
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default DreamFoundries/Qwen3.8-27B-8bit
hermes
MLX conversion of Qwen/Qwen3.8-27B, quantized with mlx-lm 0.31.3 using affine 8-bit weights and group size 64 (8.501 effective bits per weight). The safetensors weights occupy approximately 27 GB.
Comparative quality and performance benchmarks are not available for this conversion.
from mlx_lm import load, generate
model, tokenizer = load("DreamFoundries/Qwen3.8-27B-8bit")
response = generate(model, tokenizer, prompt="hello", verbose=True)
8-bit
Base model
Qwen/Qwen3.8-27B