Artex Coder 7B · MLX 4-bit

4-bit MLX build of Artex Coder 7B for Apple Silicon. Artex is a concise, accurate, to-the-point coding assistant fine-tuned from Qwen2.5-Coder-7B-Instruct. It answers in the user's language.

Usage

pip install mlx-lm
mlx_lm.generate --model 0XARTEX/artex-coder-7b-mlx-4bit --prompt "Write a Python function that checks if a string is a palindrome."
from mlx_lm import load, generate

model, tokenizer = load("0XARTEX/artex-coder-7b-mlx-4bit")
messages = [{"role": "user", "content": "How do I reverse a list in JavaScript?"}]
prompt = tokenizer.apply_chat_template(messages, add_generation_prompt=True, tokenize=False)
print(generate(model, tokenizer, prompt=prompt, max_tokens=512))

Build

Source 0XARTEX/artex-coder-7b (bf16)
Quantization 4-bit affine, group size 64 (4.5 bits per weight)
Scales / biases dtype float16
Size 4.0 GB

Scales are stored as float16, not bfloat16. Apple M1 has no hardware support for bfloat16, and an earlier internal build with bfloat16 scales decoded about 19% slower on the same file size.

Speed (measured)

Apple M1 Max, 64 GB, mlx-lm, single stream, three coding prompts, 400 max tokens:

Build Decode
bfloat16 scales (internal, unreleased) ~53 tok/s
float16 scales (this repo) ~63 tok/s

Peak memory for one loaded model is about 4.4 GB. Your numbers will vary with chip, prompt and context length.

Limitations

  • No standardized benchmark (HumanEval, MBPP) has been run yet.
  • Like any LLM, Artex can produce incorrect or insecure code. Review its output before using it.

License

Apache 2.0, the same as the base model.

Downloads last month
-
Safetensors
Model size
8B params
Tensor type
U32
·
F16
·
MLX
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for 0XARTEX/artex-coder-7b-mlx-4bit

Base model

Qwen/Qwen2.5-7B
Quantized
(1)
this model