Qwen3.6-35B-A3B KMP Dev (pruned)

A pruned Qwen3.6-35B-A3B model optimized for Kotlin Multiplatform development.

License: Apache 2.0. Based on Qwen3.6-35B-A3B by Alibaba Cloud (Apache 2.0). Modified via expert pruning.

What's kept

  • Kotlin, Swift, Gradle, Coroutines, RxSwift, Jetpack Compose, Compose Multiplatform, SwiftUI, Decompose, Metro, Koin, Ktor, Room, Coil
  • English and Russian
  • Reasoning
  • General knowledge (humanities)

What's removed

  • Other programming languages (Python, Java, Go, Rust, C, JS, PHP, Ruby, TS, HTML, bash, Qt)
  • Science/tech/esoteric domains (medicine, law, biology, chemistry, astronomy, physics, esoterics, cooking, dietetics)
  • All languages except English and Russian

Pruning method

  • Smart pruning: kept top-110 most active experts per layer (out of 256) based on heat data
  • Experts that don't activate on keep-texts are physically removed
  • Quantization: Q4_K_M (9.5GB)

Variants

Variant Format Size Quality
GGUF Q4_K_M GGUF 9.5GB ~90% of original
MLX 4-bit safetensors ~8GB ~90% of original

Usage

  • LM Studio: lms get siendsi/qwen3-6-kmp-dev
  • llama.cpp: llama-server -m Qwen3.6-35B-A3B-UD-Q4_K_M.gguf -c 16384
  • MLX: mlx_lm.generate --model siendsi/Qwen3-6-KMP-Dev-MLX-4bit

Architecture

  • 40 layers, hybrid: attention (every 4th) + SSM (Gated DeltaNet) + MoE
  • 110 experts/layer, 8 active, 1 shared
  • embedding 2048, 16 heads, 2 KV heads, head_dim 256
  • rope dim 64, mrope [11,11,10,0], freq_base 10M
Downloads last month
-
GGUF
Model size
16B params
Architecture
qwen35moe
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for siendsi/Qwen3-6-KMP-Dev-GGUF

Quantized
(736)
this model