Instructions to use Irfanuruchi/Polaris-V1-MLX-4bit with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use Irfanuruchi/Polaris-V1-MLX-4bit with MLX:
# Make sure mlx-lm is installed # pip install --upgrade mlx-lm # Generate text with mlx-lm from mlx_lm import load, generate model, tokenizer = load("Irfanuruchi/Polaris-V1-MLX-4bit") prompt = "Write a story about Einstein" messages = [{"role": "user", "content": prompt}] prompt = tokenizer.apply_chat_template( messages, add_generation_prompt=True ) text = generate(model, tokenizer, prompt=prompt, verbose=True) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Pi
How to use Irfanuruchi/Polaris-V1-MLX-4bit with Pi:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "Irfanuruchi/Polaris-V1-MLX-4bit"
Configure the model in Pi
# Install Pi: npm install -g @mariozechner/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "mlx-lm": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "Irfanuruchi/Polaris-V1-MLX-4bit" } ] } } }Run Pi
# Start Pi in your project directory: pi
- MLX LM
How to use Irfanuruchi/Polaris-V1-MLX-4bit with MLX LM:
Generate or start a chat session
# Install MLX LM uv tool install mlx-lm # Interactive chat REPL mlx_lm.chat --model "Irfanuruchi/Polaris-V1-MLX-4bit"
Run an OpenAI-compatible server
# Install MLX LM uv tool install mlx-lm # Start the server mlx_lm.server --model "Irfanuruchi/Polaris-V1-MLX-4bit" # Calling the OpenAI-compatible server with curl curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Irfanuruchi/Polaris-V1-MLX-4bit", "messages": [ {"role": "user", "content": "Hello"} ] }' - Hermes Agent
How to use Irfanuruchi/Polaris-V1-MLX-4bit with Hermes Agent:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "Irfanuruchi/Polaris-V1-MLX-4bit"
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default Irfanuruchi/Polaris-V1-MLX-4bit
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use Irfanuruchi/Polaris-V1-MLX-4bit with OpenClaw:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "Irfanuruchi/Polaris-V1-MLX-4bit"
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "Irfanuruchi/Polaris-V1-MLX-4bit" \ --custom-provider-id mlx-lm \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Polaris-V1 MLX 4-bit
This is an MLX 4-bit conversion of
nitrai-research/Polaris-V1
for inference on Apple silicon.
Quantization
- Format: MLX
- Quantization mode: affine
- Bits: 4
- Group size: 64
- Effective size: 4.503 bits per weight
- Output size: approximately 2.2 GB
Compatibility repairs
The source checkpoint identifies its text architecture as qwen3_5_text.
MLX-LM exposes the compatible text implementation as qwen3_5, so the
converted configuration was normalized accordingly.
The source configuration also used <|endoftext|> (248044) as EOS while
the tokenizer and chat template terminate responses with <|im_end|>
(248046). Both config.json and generation_config.json were corrected
to use 248046, preventing repeated stop-token generation.
Installation
pip install -U mlx-lm
Python usage
from mlx_lm import load, generate
model, tokenizer = load("Irfanuruchi/Polaris-V1-MLX-4bit")
response = generate(
model,
tokenizer,
prompt="Explain virtual memory in three concise points.",
max_tokens=256,
)
print(response)
Command-line usage
mlx_lm.generate \
--model Irfanuruchi/Polaris-V1-MLX-4bit \
--prompt "Explain virtual memory in three concise points." \
--max-tokens 256
Validation
Validated locally on an Apple M3 Pro MacBook Pro using:
- Python 3.12.14
- MLX 0.32.1
- MLX-LM 0.31.3
- Generation speed: 47.79 tokens/second
- Peak unified memory: 2.53 GB
- EOS termination: passed
- Coherent text generation: passed
Benchmark results in the metadata above are inherited from the source model card and were not independently reproduced for this quantized conversion.
License
Apache 2.0, inherited from the source model.
- Downloads last month
- 15
4-bit
Model tree for Irfanuruchi/Polaris-V1-MLX-4bit
Evaluation results
- Pass@1 on SWE-bench Verifiedself-reported31.400
- Task Completion Rate on WildClawBenchself-reported38.500
- Pass@1 on DeepSWEself-reported26.800