Interim Labs Qwen3.5 model series

Interim Labs Qwen3.5-4B · MLX 4-bit

A standalone text-only MLX 4-bit release published by Interim Labs for local text generation on Apple Silicon. This package derives from Qwen/Qwen3.5-4B. Interim Labs packages and documents this derivative; the upstream model and conversion sources are credited below.

Model series

Part of the Interim Labs Qwen3.5 series: three sizes, each with an original Qwen3.5 variant and a Huihui abliterated variant, in two formats.

Model MLX 4-bit GGUF Q4_K_M
Qwen3.5 2B MLX GGUF
Huihui Qwen3.5 2B abliterated MLX GGUF
Qwen3.5 4B MLX GGUF
Huihui Qwen3.5 4B abliterated MLX GGUF
Qwen3.5 9B MLX GGUF
Huihui Qwen3.5 9B abliterated MLX GGUF

Package details

Property Value
Parameter class 4B
Format MLX / Safetensors
Quantization 4-bit affine, group size 64
Weight file model.safetensors
Weight size 2.37 GB (2,367,224,773 bytes)
Model type qwen3_5_text
Configured context 8,192 tokens
Retained text tensors 924

This package contains language-model tensors only and does not provide image input. Text-only derivation removed 297 vision tensors. The packaged configuration sets the context to 8,192 tokens; this is not the upstream model’s full advertised context length.

Usage and runtime requirements

Download the package with the Hugging Face CLI:

hf download interimlabs/InterimLabs-Qwen3.5-4B-MLX-4bit

Use an MLX runtime on Apple Silicon that supports the qwen3_5_text architecture and this package’s quantization. Use the included tokenizer and chat template, and keep the total context within 8,192 tokens. Generic Python mlx_lm builds without qwen3_5_text support may fail to load this package. Runtime compatibility across applications has not been comprehensively verified.

Sources and conversion

The conversion source does not establish the exact upstream revision used for its conversion. The reference revision above documents the inspected upstream snapshot, not a verified conversion lineage.

See provenance.json for source and package metadata.

Weight-file SHA-256:

799b8fe437040924178c0ce28acb4fb0a5830a4a6f0ddcff004e13ab468bfbf8

Evaluation and limitations

This release does not include a comprehensive benchmark of the packaged model. No writing-quality, reliability, or performance improvement over the credited source is claimed.

License and credits

Distributed under Apache-2.0. Credit belongs to the Qwen team and the conversion authors linked above. Interim Labs publishes this package and its documentation. See the linked upstream cards for the original model details and limitations.

Downloads last month
26
Safetensors
Model size
4B params
Tensor type
U32
·
BF16
·
F32
·
MLX
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for interimlabs/InterimLabs-Qwen3.5-4B-MLX-4bit

Finetuned
Qwen/Qwen3.5-4B
Quantized
(512)
this model

Collection including interimlabs/InterimLabs-Qwen3.5-4B-MLX-4bit