Agnes-3.0-Flash — MLX 4-bit

4-bit MLX quantization of Agnes-AI/Agnes-3.0-Flash (Apache-2.0) — loads as a stock qwen3_5 model with no custom code.

  • Affine 4-bit, group size 64 — 4.50 bits/weight, 17 GB
  • 262,144-token context, thinking on/off via the original chat template
  • Text only: MTP head and vision tower not included

What changed

Converted from the original Agnes format to standard Qwen3.5 architecture:

  • Folded parallel FFN into main MLP via concatenation (intermediate_size: 19456)
  • Renamed delta_attnlinear_attn, global_attnself_attn
  • Converted one-centered RMSNorm to standard format
  • Cast bf16 → fp16 for serialization compatibility
  • Stripped MTP weights (prevents double-conversion in mlx_lm)

Usage

pip install mlx-lm
mlx_lm.generate --model hermitdave/Agnes-3.0-Flash-MLX-4bit --prompt "Hello" --max-tokens 200

Drop the folder under ~/.lmstudio/models/hermitdave/ in LM Studio and it appears as a qwen3_5 model.

Speculative decoding with MTP drafter

A companion MTP drafter is available for speculative decoding (up to 2× faster generation):

pip install mlx-vlm
python -m mlx_vlm.server \
    --model hermitdave/Agnes-3.0-Flash-MLX-4bit \
    --draft-model hermitdave/Agnes-3.0-Flash-MTP-drafter

Or with the Python API:

from mlx_lm import load
from mlx_vlm.speculative.drafters.qwen3_5_mtp.config import Qwen3_5MTPConfig
from mlx_vlm.speculative.drafters.qwen3_5_mtp.qwen3_5_mtp import Qwen3_5MTPDraftModel
import json, mlx.core as mx, safetensors.torch

base_model, tokenizer = load("hermitdave/Agnes-3.0-Flash-MLX-4bit")
config = Qwen3_5MTPConfig.from_dict(json.load(open("path/to/Agnes-3.0-Flash-MTP-drafter/config.json")))
mtp = Qwen3_5MTPDraftModel(config)
weights = safetensors.torch.load_file("path/to/Agnes-3.0-Flash-MTP-drafter/model.safetensors")
mtp.load_weights([(k, mx.array(v)) for k, v in weights.items()])
mtp.bind(base_model)

The drafter is architecture-compatible with Qwen3.5's gated attention and was extracted from the original Agnes-3.0-Flash model. See the MTP drafter repo for details.

Note: oMLX does not yet support MTP for text-only models. Use mlx_vlm.server or the Python API.

Attribution

This conversion was produced by Hermes Agent (Nous Research) — the autonomous research and conversion pipeline that identified the correct quantization parameters, fixed one-centered norm conversion, and validated output quality. Verified against the reference verison/Agnes-3.0-Flash-MLX-4bit model.

Downloads last month
43
Safetensors
Model size
32B params
Tensor type
U32
·
BF16
·
F16
·
MLX
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for hermitdave/Agnes-3.0-Flash-MLX-4bit

Quantized
(12)
this model