KAT-Coder-V2.5-Dev-MLX-MXFP4

MLX MXFP4 (mxfp4, group size 32) quantized variant of Kwaipilot/KAT-Coder-V2.5-Dev for Apple silicon via mlx-lm.

Provenance

  • Source: Kwaipilot/KAT-Coder-V2.5-Dev @ revision 7be56fe773e72b6f5ca93c1ae45d828ddb893922 (Apache-2.0).
  • Quantized with mlx_lm.convert (mlx-lm 0.31.3): mxfp4, 4-bit, group size 32.
  • Text-only pack: the upstream checkpoint is multimodal (Qwen3_5MoeForConditionalGeneration); mlx-lm's qwen3_5_moe loader drops the vision tower (model.visual.*) by design, so this pack ships only the text MoE. Use the upstream repo if you need vision.

Smoke gate

Before upload this pack passed a deterministic coherence gate: greedy 64-token chat generation loaded through mlx_lm.load, judged for emptiness, repetition loops, multi-script gibberish, and special-token debris. Verdict: ok.

Usage

pip install mlx-lm
mlx_lm.generate --model majentik/KAT-Coder-V2.5-Dev-MLX-MXFP4 --prompt "Write a binary search in Python"

Evaluation

Benchmark Score
arc_easy_acc 0.7350
hellaswag_acc 0.5250

Available tiers

Downloads last month
551
Safetensors
Model size
7B params
Tensor type
U8
U32
BF16
MLX
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 馃檵 Ask for provider support

Model tree for majentik/KAT-Coder-V2.5-Dev-MLX-MXFP4

Quantized
(66)
this model