Qwen3.5-2B OneCompression 4-bit MLX

Text-only MLX checkpoint produced from Qwen/Qwen3.5-2B for AyaneSDK's on-device conversation example.

  • OneCompression GPTQ with quantization-error propagation (QEP)
  • 4-bit weights, group size 128
  • 256 Japanese dialogue calibration samples of 512 tokens
  • 186 quantized linear layers
  • Token embedding quantized separately to asymmetric MLX 4-bit
  • Vision weights are not included

The packed GPTQ-v1 linear weights were converted losslessly to MLX's row-major affine representation. The model configuration retains an onecompression_source_quantization audit record.

Source

Base model: Qwen/Qwen3.5-2B

Quantizer: FujitsuResearch/OneCompression

Downloads last month
-
Safetensors
Model size
0.3B params
Tensor type
F32
U32
BF16
MLX
Hardware compatibility
Log In to add your hardware

Quantized

Inference Providers NEW
This model isn't deployed by any Inference Provider. 馃檵 Ask for provider support

Model tree for kizuna-intelligence/Qwen3.5-2B-OneCompression-4bit-MLX

Finetuned
Qwen/Qwen3.5-2B
Finetuned
(335)
this model