Instructions to use interimlabs/InterimLabs-Qwen3.5-4B-MLX-4bit with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use interimlabs/InterimLabs-Qwen3.5-4B-MLX-4bit with MLX:
# Make sure mlx-lm is installed # pip install --upgrade mlx-lm # Generate text with mlx-lm from mlx_lm import load, generate model, tokenizer = load("interimlabs/InterimLabs-Qwen3.5-4B-MLX-4bit") prompt = "Write a story about Einstein" messages = [{"role": "user", "content": prompt}] prompt = tokenizer.apply_chat_template( messages, add_generation_prompt=True ) text = generate(model, tokenizer, prompt=prompt, verbose=True) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Pi
How to use interimlabs/InterimLabs-Qwen3.5-4B-MLX-4bit with Pi:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "interimlabs/InterimLabs-Qwen3.5-4B-MLX-4bit"
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "mlx-lm": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "interimlabs/InterimLabs-Qwen3.5-4B-MLX-4bit" } ] } } }Run Pi
# Start Pi in your project directory: pi
- MLX LM
How to use interimlabs/InterimLabs-Qwen3.5-4B-MLX-4bit with MLX LM:
Generate or start a chat session
# Install MLX LM uv tool install mlx-lm # Interactive chat REPL mlx_lm.chat --model "interimlabs/InterimLabs-Qwen3.5-4B-MLX-4bit"
Run an OpenAI-compatible server
# Install MLX LM uv tool install mlx-lm # Start the server mlx_lm.server --model "interimlabs/InterimLabs-Qwen3.5-4B-MLX-4bit" # Calling the OpenAI-compatible server with curl curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "interimlabs/InterimLabs-Qwen3.5-4B-MLX-4bit", "messages": [ {"role": "user", "content": "Hello"} ] }' - Hermes Agent
How to use interimlabs/InterimLabs-Qwen3.5-4B-MLX-4bit with Hermes Agent:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "interimlabs/InterimLabs-Qwen3.5-4B-MLX-4bit"
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default interimlabs/InterimLabs-Qwen3.5-4B-MLX-4bit
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use interimlabs/InterimLabs-Qwen3.5-4B-MLX-4bit with OpenClaw:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "interimlabs/InterimLabs-Qwen3.5-4B-MLX-4bit"
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "interimlabs/InterimLabs-Qwen3.5-4B-MLX-4bit" \ --custom-provider-id mlx-lm \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Interim Labs Qwen3.5-4B · MLX 4-bit
A standalone text-only MLX 4-bit release published by Interim Labs for local text generation on Apple Silicon. This package derives from Qwen/Qwen3.5-4B. Interim Labs packages and documents this derivative; the upstream model and conversion sources are credited below.
Model series
Part of the Interim Labs Qwen3.5 series: three sizes, each with an original Qwen3.5 variant and a Huihui abliterated variant, in two formats.
| Model | MLX 4-bit | GGUF Q4_K_M |
|---|---|---|
| Qwen3.5 2B | MLX | GGUF |
| Huihui Qwen3.5 2B abliterated | MLX | GGUF |
| Qwen3.5 4B | MLX | GGUF |
| Huihui Qwen3.5 4B abliterated | MLX | GGUF |
| Qwen3.5 9B | MLX | GGUF |
| Huihui Qwen3.5 9B abliterated | MLX | GGUF |
Package details
| Property | Value |
|---|---|
| Parameter class | 4B |
| Format | MLX / Safetensors |
| Quantization | 4-bit affine, group size 64 |
| Weight file | model.safetensors |
| Weight size | 2.37 GB (2,367,224,773 bytes) |
| Model type | qwen3_5_text |
| Configured context | 8,192 tokens |
| Retained text tensors | 924 |
This package contains language-model tensors only and does not provide image input. Text-only derivation removed 297 vision tensors. The packaged configuration sets the context to 8,192 tokens; this is not the upstream model’s full advertised context length.
Usage and runtime requirements
Download the package with the Hugging Face CLI:
hf download interimlabs/InterimLabs-Qwen3.5-4B-MLX-4bit
Use an MLX runtime on Apple Silicon that supports the qwen3_5_text architecture and this package’s quantization. Use the included tokenizer and chat template, and keep the total context within 8,192 tokens. Generic Python mlx_lm builds without qwen3_5_text support may fail to load this package. Runtime compatibility across applications has not been comprehensively verified.
Sources and conversion
- MLX source: mlx-community/Qwen3.5-4B-4bit at
0e7ffd5c629ef7719d4cbc04069232580bfa9d9c. - Reference upstream: Qwen/Qwen3.5-4B at
851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a.
The conversion source does not establish the exact upstream revision used for its conversion. The reference revision above documents the inspected upstream snapshot, not a verified conversion lineage.
See provenance.json for source and package metadata.
Weight-file SHA-256:
799b8fe437040924178c0ce28acb4fb0a5830a4a6f0ddcff004e13ab468bfbf8
Evaluation and limitations
This release does not include a comprehensive benchmark of the packaged model. No writing-quality, reliability, or performance improvement over the credited source is claimed.
License and credits
Distributed under Apache-2.0. Credit belongs to the Qwen team and the conversion authors linked above. Interim Labs publishes this package and its documentation. See the linked upstream cards for the original model details and limitations.
- Downloads last month
- 26
4-bit
