Qwen3.5-2B AutoRound A16W4 for Modalix

Runtime artifacts compiled for the LLiMa runtime on SiMa.ai Modalix.

  • SmoothQuant alpha: 0.5.
  • Decoder Linear weights, including DeltaNet QKV/Z/output projections: symmetric AutoRound INT4, group size 256.
  • Output head: symmetric GPTQ INT4, group size 256.
  • Vision and projector Linear weights: RTN INT8 per output channel.
  • DeltaNet A/B projections, convolution parameters, normalization parameters, A_log, and dt_bias: BF16.
  • Context capacity: 4096 tokens; prefill group size: 128.
  • Vision input: 448 × 448; vision encoder packaged as per-layer ELFs.
  • Filter sharing, quantized embeddings, and quantized KV cache enabled.

Run

Download the complete repository while preserving its directory structure. On Modalix with a compatible LLiMa runtime installed:

llima run /path/to/Qwen3.5-2B-Autoround-a16w4-Modalix

devkit/ contains runtime configuration, tokenizer assets, and embeddings. elf_files/ contains 119 compiled MLA programs. This repository contains Modalix runtime artifacts rather than a Transformers checkpoint.

Compilation, archive validation, and local llima-deploy completed successfully. Modalix runtime evaluation is pending.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for florianvoss/Qwen3.5-2B-Autoround-a16w4-Modalix

Finetuned
Qwen/Qwen3.5-2B
Finetuned
(370)
this model

Collection including florianvoss/Qwen3.5-2B-Autoround-a16w4-Modalix