AX-Tiel-Coder-35B-A3B-MLX-AXQ-MXFP4-MTP

Development repack. Not certified. No measured quality-parity or MTP-speed claim.

This is an AXQuant MXFP4 requantization of the exact peculiar-ragdoll/Tiel-Coder-35B-A3B-MLX-oQ6e-MTP checkpoint, revision 88625754ac91b542280a5602239ce6b2166366f0. The source is already mixed-precision oQ6e, not original BF16. Public MLX APIs dequantize its weights and repack the language trunk using an AXQuant manual recipe. This adds quantization error and does not restore the original full-precision weights. Upstream benchmark results and coding-performance claims do not apply to this repack. No upstream importance matrix, calibration corpus, or quantizer implementation is imported.

Artifact

Property Value
AXQuant version 1.9.0
Architecture Qwen3.5-class 35B-A3B MoE; qwen35-moe-v1
Language trunk Native MLX MXFP4, group size 32
Protection floors Embeddings/routers at least 8-bit; norms/LM head BF16
Measured main BPW 4.634223
Measured total BPW 4.901270
Weight bytes 22026200211
MTP 785 source tensors preserved in mtp.safetensors
Vision Source vision tensors preserved in vision.safetensors
Certification None; architecture-prior manual allocation

The protected MTP and vision tensor payloads were compared byte-for-byte with the pinned source. The source tokenizer and Sharp chat template are preserved. Basic deterministic text-generation smoke tests ran with MLX-LM on the factory host. These checks establish conversion integrity and basic loadability, not model quality, vision capability, MTP acceptance, or acceleration. AX Engine execution was not measured for this artifact in this campaign.

Download and basic text generation

python -m pip install mlx-lm huggingface_hub
hf download AutomatosX/AX-Tiel-Coder-35B-A3B-MLX-AXQ-MXFP4-MTP --local-dir ./AX-Tiel-Coder-35B-A3B-MLX-AXQ-MXFP4-MTP
mlx_lm.generate --model ./AX-Tiel-Coder-35B-A3B-MLX-AXQ-MXFP4-MTP --max-tokens 256 --prompt "Write a Python add function."

Use an Apple Silicon Mac with sufficient unified memory. Pin the published Hub commit for reproducible deployments. The tested packaging environment used MLX 0.32.2 and MLX-LM 0.31.3. MXFP4 here is the native MLX Apple format, not NVIDIA NVFP4.

MTP and multimodal scope

The -MTP suffix means the head is packaged, not that acceleration is certified or enabled by the text command above. Stock MLX-LM text generation does not use the separate MTP sidecar. A compatible sidecar-aware runtime is required, and runtime-specific MTP enablement must be validated separately. No MTP grafting or new training was performed. The upstream MTP provenance remains that of the exact source repository. Vision weights are retained but this campaign does not claim vision inference support.

Provenance and license

The upstream model card declares MIT. Follow its license terms and usage information. Credits: peculiar-ragdoll for the exact source build and template, Ornith for the model lineage, and the additional upstream contributors identified in the source model card. AXQuant uses public MLX conversion APIs. See axquant_plan.json, axquant_quantizer_execution.json, the protected-sidecar manifests, and axquant_manifest.json for allocation and file bindings.

Downloads last month
309
Safetensors
Model size
35B params
Tensor type
BF16
·
U32
·
MLX
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for AutomatosX/AX-Tiel-Coder-35B-A3B-MLX-AXQ-MXFP4-MTP