basal-1.5-mini-MLX-bf16

bf16 MLX copy (no quantization) of Remek/basal-1.5-mini (1.5B) for Apple Silicon, by Remigiusz Kinas. Same model, same prompt format and the same typed answers; see the main card and basal.si5.pl for the description, benchmarks, limitations and license. The official MLX ports are -MLX-8bit (recommended) and -MLX-fp4; this repo adds a variant they do not cover. Community conversion, not an official release.

Equivalent to the bf16 model on this check: the same decision on all 404 items, max |dp| 0.038; the fastest MLX port, at twice the memory of -MLX-6bit.

Converting basal-1.5 to MLX yourself? Write rope_theta into the config. The basal-1.5 configs keep rope_theta (1,000,000) inside the transformers-5 rope_parameters field, and mlx_lm's Llama (checked: 0.31.3 and 0.32.0) falls back to 10,000 without it, as Remek documents in basal's HARDWARE.md. basal-serve --mode mlx corrects this itself, but anything else that loads the converted weights through mlx-lm (mlx_lm.server, the Python API, vllm-metal) answers without a warning, and wrongly: in our raw mlx-lm readout of 40 Polish decisions, agreement with the PyTorch bf16 model was 0.975 with probability shifts up to 0.32 (with the value set: 1.000 and 0.038). This repo and Remek's own MLX ports have "rope_theta": 1000000 at the top level of config.json. Conversion script: convert_mlx.py.

Check against the bf16 model

404 Polish decisions from two independent sets (experiments 03 and 04 of system-1-ai: customer service, and 12 domains from banking to HR; yes/no, choice and score questions), both option orders, calibrated, compared with the same weights in PyTorch bf16 (accuracy 0.847). Apple M5 Max. Full write-up, all ports side by side and per-category tables: quant_results.md.

this port
agreement with PyTorch bf16 (same top option) 1.000
decisions that differ / of them where bf16 had a margin ≥ 0.2 0 / 0
accuracy (PyTorch bf16: 0.847) 0.847
median latency per decision (e03 / e04) 42 / 45 ms
weights 3.2 GB

404 items is a check of the port, not a benchmark: accuracy differences of 2 points or less are noise; agreement and the differing decisions are the signal.

Weights: all weights in bf16, as in the original model. Converted with mlx_lm 0.31.3; rope_theta (1e6) is written at the top level of config.json, because mlx_lm does not read the transformers-5 rope_parameters field of the base config and would silently fall back to 10000 (a plain conversion loads and answers, but wrongly).

Files: the MLX weights and config, tokenizer and chat template, CALIBRATION.json, basal.json (copied unchanged).

Quick start

uv venv --python 3.12 ~/basal-mlx && source ~/basal-mlx/bin/activate
uv pip install "basal[mlx] @ https://github.com/rkinas/basal/archive/refs/tags/v1.5.0.tar.gz"
basal-serve --model agentGreg/basal-1.5-mini-MLX-bf16 --mode mlx --port 8000
curl -s localhost:8000/v1/systemone -H 'content-type: application/json' -d '{
  "state": "Klient: od wczoraj nie mogę zalogować się do bankowości internetowej, system pokazuje błąd hasła.",
  "questions": {"dept": {"type": "choice", "instructions": "Do którego działu skierować zgłoszenie?",
    "criteria": {"cards": "Reklamacje kart", "online": "Wsparcie bankowości elektronicznej", "loans": "Kredyty"}}}}'

Evidence spans are not available with MLX (use the bf16 model with the PyTorch engine).

Calibration

Temperatures need no change for this port (label-free paired fit: t = 1.000 for yes/no, 1.005 for choice), so it ships only the bf16 model's CALIBRATION.json. The confidence thresholds are the bf16 model's and are not certified for this port: refit them on your own labelled requests before automating decisions with them.

License

Apache-2.0, like basal-1.5-mini.

Downloads last month
324
Safetensors
Model size
2B params
Tensor type
BF16
·
MLX
Hardware compatibility
Log In to add your hardware

Quantized

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for agentGreg/basal-1.5-mini-MLX-bf16

Finetuned
(1)
this model