kas-4b

A fine-tuned decision engine based on Qwen/Qwen3-4B, submitted to the Decision Index v0.2.1 leaderboard.

Decision Index: 40.1 (raw 54.18) โ€” above Tev1-4B (29.2), within 6 points of Scion v4 9B (45.9).

Scores

Area Score
Tools & Automation 49.0
Retrieval & Classification 40.4
Knowledge & Reasoning 39.5
Arts & Human Taste 37.7
Language Understanding 35.2
Decision Index 40.1

Full results: kpiya/decision-index-results


4B Model Comparison โ€” Where kas-4b Leads and Trails

Per-area breakdown against all 4B submissions on Decision Index 0.2.1.

Area ezjev-4b (51.2) Nox 4B (43.8) kas-4b (40.1) intelif-4B (31.8)
Tools & Automation 69.9 60.1 49.0 51.0
Retrieval & Classification 56.3 52.4 40.4 39.6
Knowledge & Reasoning 33.5 27.6 39.5 18.3
Language Understanding 60.2 48.6 35.2 31.1
Arts & Human Taste 28.8 25.8 37.7 17.8

kas-4b leads the 4B tier on:

  • Knowledge & Reasoning (39.5) โ€” highest of all 4B models; higher LoRA rank (r64) likely helps on harder reasoning tasks
  • Arts & Human Taste (37.7) โ€” highest of all 4B models by a wide margin

kas-4b trails on:

  • Language Understanding (35.2) โ€” 13โ€“25 points behind ezjev and Nox; language-heavy fine-tuning data favors those models
  • Tools & Automation (49.0) โ€” behind ezjev (69.9) and Nox (60.1)

Profile: the most balanced 4B model on the benchmark โ€” consistent across all five areas rather than peaking on one or two.


Training

  • Base model: Qwen/Qwen3-4B (Apache 2.0, pure transformer)
  • Method: LoRA fine-tune (PEFT), adapter merged into base weights
  • LoRA config: rank 64, alpha 128, 2 epochs, LR 1e-4, batch size 16, 4096 token limit
  • Recipe: inspired by Scion (Sinan Ozdemir)
  • Hardware: NVIDIA RTX PRO 6000 Blackwell (HF Jobs)

Engine

This model is served via a custom engine (kas_engine.engine:KasEngine) that:

  • Maps options to single-token uppercase labels (Aโ€“Z, then two-letter)
  • Repeats the JSON prompt once (Scion-style)
  • Scores label logits at the generation position, softmaxed to probabilities
  • Refuses over-context prompts โ€” never truncates

Engine code: kashavpiya/kas-4b

Latency

Measured on RTX PRO 6000, single process, one request at a time:

Metric Value
Median 39.2 ms
p95 546.4 ms
Mean 153.3 ms

Evaluation

Run with Decision Index kit v0.2.1 on 150,759 rows (184 unsupported / over-context, not truncated).

Notes

Not trained on Decision Index suite data. Base model license: Apache 2.0.

Downloads last month
22
Safetensors
Model size
4B params
Tensor type
BF16
ยท
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for kpiya/kas-4b

Finetuned
Qwen/Qwen3-4B
Adapter
(1184)
this model