Qwen3.5-4B INT8 โ€” LLiMa runtime package

Compiled model artifacts for the LLiMa runtime on SiMa.ai Modalix.

  • BF16 activations and INT8 weights (A_BF16_W_INT8).
  • Maximum token count: 4096; prefill group size: 128.
  • Compiled vision input: 32 ร— 32.
  • Filter sharing, quantized embeddings, and quantized KV cache enabled.

Run

Copy or download the entire repository, preserving the directory structure. On a Modalix device with a compatible LLiMa runtime installed:

llima run /path/to/Qwen3.5-4B-INT8

devkit/ contains model configuration, tokenizer files, and embeddings. elf_files/ contains the compiled MLA programs. These are device runtime artifacts; they are not a Transformers checkpoint.

Compilation and local deployment completed. This package has not been validated by running the full model on a Modalix device.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for florianvoss/Qwen3.5-4B-INT8-Modalix

Finetuned
Qwen/Qwen3.5-4B
Finetuned
(608)
this model

Collection including florianvoss/Qwen3.5-4B-INT8-Modalix