Instructions to use 0XARTEX/artex-coder-7b-mlx-4bit with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use 0XARTEX/artex-coder-7b-mlx-4bit with MLX:
# Make sure mlx-lm is installed # pip install --upgrade mlx-lm # Generate text with mlx-lm from mlx_lm import load, generate model, tokenizer = load("0XARTEX/artex-coder-7b-mlx-4bit") prompt = "Write a story about Einstein" messages = [{"role": "user", "content": prompt}] prompt = tokenizer.apply_chat_template( messages, add_generation_prompt=True ) text = generate(model, tokenizer, prompt=prompt, verbose=True) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Pi
How to use 0XARTEX/artex-coder-7b-mlx-4bit with Pi:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "0XARTEX/artex-coder-7b-mlx-4bit"
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "mlx-lm": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "0XARTEX/artex-coder-7b-mlx-4bit" } ] } } }Run Pi
# Start Pi in your project directory: pi
- MLX LM
How to use 0XARTEX/artex-coder-7b-mlx-4bit with MLX LM:
Generate or start a chat session
# Install MLX LM uv tool install mlx-lm # Interactive chat REPL mlx_lm.chat --model "0XARTEX/artex-coder-7b-mlx-4bit"
Run an OpenAI-compatible server
# Install MLX LM uv tool install mlx-lm # Start the server mlx_lm.server --model "0XARTEX/artex-coder-7b-mlx-4bit" # Calling the OpenAI-compatible server with curl curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "0XARTEX/artex-coder-7b-mlx-4bit", "messages": [ {"role": "user", "content": "Hello"} ] }' - Hermes Agent
How to use 0XARTEX/artex-coder-7b-mlx-4bit with Hermes Agent:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "0XARTEX/artex-coder-7b-mlx-4bit"
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default 0XARTEX/artex-coder-7b-mlx-4bit
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use 0XARTEX/artex-coder-7b-mlx-4bit with OpenClaw:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "0XARTEX/artex-coder-7b-mlx-4bit"
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "0XARTEX/artex-coder-7b-mlx-4bit" \ --custom-provider-id mlx-lm \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Artex Coder 7B · MLX 4-bit
4-bit MLX build of Artex Coder 7B for Apple Silicon. Artex is a concise, accurate, to-the-point coding assistant fine-tuned from Qwen2.5-Coder-7B-Instruct. It answers in the user's language.
Usage
pip install mlx-lm
mlx_lm.generate --model 0XARTEX/artex-coder-7b-mlx-4bit --prompt "Write a Python function that checks if a string is a palindrome."
from mlx_lm import load, generate
model, tokenizer = load("0XARTEX/artex-coder-7b-mlx-4bit")
messages = [{"role": "user", "content": "How do I reverse a list in JavaScript?"}]
prompt = tokenizer.apply_chat_template(messages, add_generation_prompt=True, tokenize=False)
print(generate(model, tokenizer, prompt=prompt, max_tokens=512))
Build
| Source | 0XARTEX/artex-coder-7b (bf16) |
| Quantization | 4-bit affine, group size 64 (4.5 bits per weight) |
| Scales / biases dtype | float16 |
| Size | 4.0 GB |
Scales are stored as float16, not bfloat16. Apple M1 has no hardware support for bfloat16, and an earlier internal build with bfloat16 scales decoded about 19% slower on the same file size.
Speed (measured)
Apple M1 Max, 64 GB, mlx-lm, single stream, three coding prompts, 400 max tokens:
| Build | Decode |
|---|---|
| bfloat16 scales (internal, unreleased) | ~53 tok/s |
| float16 scales (this repo) | ~63 tok/s |
Peak memory for one loaded model is about 4.4 GB. Your numbers will vary with chip, prompt and context length.
Limitations
- No standardized benchmark (HumanEval, MBPP) has been run yet.
- Like any LLM, Artex can produce incorrect or insecure code. Review its output before using it.
License
Apache 2.0, the same as the base model.
- Downloads last month
- -
4-bit
Model tree for 0XARTEX/artex-coder-7b-mlx-4bit
Base model
Qwen/Qwen2.5-7B