Occamy 1.0 — MLX 3-bit, group size 64

Candidate release. Mac Metal acceptance is pending. Linux native MLX validation passed; this is not a claim of validated Mac performance or full model quality.

Source: Accio-Lab/occamy-1.0, revision 8f8e0e58a3c9df042be1a3fa2c191fd8047acfb8.

Converted using official mlx 0.32.2 and mlx-lm 0.31.3 native affine quantization. Base precision is 3-bit with group size 64; native router and shared-expert gate modules use 8-bit. This is a text-only qwen3_5_moe export; vision and MTP are not included.

A lossless adapter stacks separate expert weights in numeric expert order before invoking the Qwen3.5 sanitizer exactly once. Quantization and serialization use native APIs. Reload uses the stock loader without an adapter.

Validation: strict stock reload, complete stored floating-value checks, native dequantization of every quantized row, tokenizer/template comparison, and one bounded cached greedy CPU generation with finite logits. The prompt “Compute 2+2. Answer briefly.” returned 4. This limited smoke test is not a quality benchmark.

from mlx_lm import load, generate
model, tokenizer = load("Accio-Lab/occamy-1.0-MLX-3bit")

See SHA256SUMS for artifact hashes and validation_summary.json for validation scope.

Downloads last month
95
Safetensors
Model size
35B params
Tensor type
U32
·
BF16
·
MLX
Hardware compatibility
Log In to add your hardware

3-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Accio-Lab/occamy-1.0-MLX-3bit

Quantized
(17)
this model