Instructions to use mlx-community/KAT-Coder-V2.5-Dev-OptiQ-4bit-REAP-18B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use mlx-community/KAT-Coder-V2.5-Dev-OptiQ-4bit-REAP-18B with MLX:
# Make sure mlx-lm is installed # pip install --upgrade mlx-lm # Generate text with mlx-lm from mlx_lm import load, generate model, tokenizer = load("mlx-community/KAT-Coder-V2.5-Dev-OptiQ-4bit-REAP-18B") prompt = "Write a story about Einstein" messages = [{"role": "user", "content": prompt}] prompt = tokenizer.apply_chat_template( messages, add_generation_prompt=True ) text = generate(model, tokenizer, prompt=prompt, verbose=True) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Pi
How to use mlx-community/KAT-Coder-V2.5-Dev-OptiQ-4bit-REAP-18B with Pi:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "mlx-community/KAT-Coder-V2.5-Dev-OptiQ-4bit-REAP-18B"
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "mlx-lm": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "mlx-community/KAT-Coder-V2.5-Dev-OptiQ-4bit-REAP-18B" } ] } } }Run Pi
# Start Pi in your project directory: pi
- MLX LM
How to use mlx-community/KAT-Coder-V2.5-Dev-OptiQ-4bit-REAP-18B with MLX LM:
Generate or start a chat session
# Install MLX LM uv tool install mlx-lm # Interactive chat REPL mlx_lm.chat --model "mlx-community/KAT-Coder-V2.5-Dev-OptiQ-4bit-REAP-18B"
Run an OpenAI-compatible server
# Install MLX LM uv tool install mlx-lm # Start the server mlx_lm.server --model "mlx-community/KAT-Coder-V2.5-Dev-OptiQ-4bit-REAP-18B" # Calling the OpenAI-compatible server with curl curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "mlx-community/KAT-Coder-V2.5-Dev-OptiQ-4bit-REAP-18B", "messages": [ {"role": "user", "content": "Hello"} ] }' - Hermes Agent
How to use mlx-community/KAT-Coder-V2.5-Dev-OptiQ-4bit-REAP-18B with Hermes Agent:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "mlx-community/KAT-Coder-V2.5-Dev-OptiQ-4bit-REAP-18B"
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default mlx-community/KAT-Coder-V2.5-Dev-OptiQ-4bit-REAP-18B
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use mlx-community/KAT-Coder-V2.5-Dev-OptiQ-4bit-REAP-18B with OpenClaw:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "mlx-community/KAT-Coder-V2.5-Dev-OptiQ-4bit-REAP-18B"
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "mlx-community/KAT-Coder-V2.5-Dev-OptiQ-4bit-REAP-18B" \ --custom-provider-id mlx-lm \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
mlx-community/KAT-Coder-V2.5-Dev-OptiQ-4bit-REAP-18B
Built with mlx-optiq, the MLX-native toolkit to quantize, prune, fine-tune, and serve LLMs locally on Apple Silicon. All OptiQ models · Docs
11.4 GB instead of 20.4 GB. 14.4 GB of memory to run.
| Parent | This model | ||
|---|---|---|---|
| On disk | 20.4 GB | 11.4 GB | −44% |
| Parameters | 34.7B | 18.3B | −47% |
50% of the routed experts are removed from mlx-community/KAT-Coder-V2.5-Dev-OptiQ-4bit; active parameters per token are unchanged, since top-8 routing is preserved and only the stored expert bank shrinks. That is why it gets smaller without getting slower.
Retained experts are copied bit-for-bit from the parent quant. Nothing is dequantized, re-quantized, merged, or retrained.
This variant was not separately benchmarked. It is published under the recipe validated end to end on Qwen3.6-35B-A3B-OptiQ-4bit-REAP-19B, the same architecture at the same 50 % retention: Capability Score 80.03 -> 76.57, with the loss concentrated in MMLU (-21.4) and procedural ability intact (GSM8K +2.6, IFEval +4.3, BFCL -1.0, HumanEval -1.3).
Two things were measured on this checkpoint. The ranking rule was chosen by scoring both candidates against the unpruned model, which picked the conditional mean. And the resulting divergence from the unpruned parent is KL 0.213 — for reference, the checkpoints that degrade visibly under pruning measure above 1.0, and this one is well inside the range where generation is indistinguishable in review.
Details
| Property | Value |
|---|---|
| Experts retained | 128 of 256 per layer |
| Active experts per token | 8 (unchanged) |
| Allocation | uniform (128 of 256 in every layer) |
| Size | 11.4 GB (parent 20.4 GB) |
| Parameters | 18.3B (parent 34.7B) |
| Selection | REAP — mean of router weight x expert output norm, over the tokens each expert served |
| Calibration | optiq six-domain mix, 8 samples |
| MTP sidecar | absent |
Use it
pip install mlx-optiq
optiq serve --model mlx-community/KAT-Coder-V2.5-Dev-OptiQ-4bit-REAP-18B
from mlx_lm import load, generate
model, tok = load("mlx-community/KAT-Coder-V2.5-Dev-OptiQ-4bit-REAP-18B")
print(generate(model, tok, prompt="Hello", max_tokens=64))
Method
Expert pruning follows REAP (Cerebras Research, ICLR 2026 — REAP the Experts: Why Pruning Prevails for One-Shot MoE Compression). Experts are ranked by the conditional mean of router weight × expert output norm over calibration data; the lowest-ranked are removed and the router is sliced to match.
OptiQ applies it in the quantized domain — directly on a quantized checkpoint, with no BF16 parent and no dequantization of survivors — via optiq prune-experts. See the pruning docs.
- Downloads last month
- -
4-bit
Model tree for mlx-community/KAT-Coder-V2.5-Dev-OptiQ-4bit-REAP-18B
Base model
Kwaipilot/KAT-Coder-V2.5-Dev