kas-4b
A fine-tuned decision engine based on Qwen/Qwen3-4B, submitted to the Decision Index v0.2.1 leaderboard.
Decision Index: 40.1 (raw 54.18) โ above Tev1-4B (29.2), within 6 points of Scion v4 9B (45.9).
Scores
| Area | Score |
|---|---|
| Tools & Automation | 49.0 |
| Retrieval & Classification | 40.4 |
| Knowledge & Reasoning | 39.5 |
| Arts & Human Taste | 37.7 |
| Language Understanding | 35.2 |
| Decision Index | 40.1 |
Full results: kpiya/decision-index-results
4B Model Comparison โ Where kas-4b Leads and Trails
Per-area breakdown against all 4B submissions on Decision Index 0.2.1.
| Area | ezjev-4b (51.2) | Nox 4B (43.8) | kas-4b (40.1) | intelif-4B (31.8) |
|---|---|---|---|---|
| Tools & Automation | 69.9 | 60.1 | 49.0 | 51.0 |
| Retrieval & Classification | 56.3 | 52.4 | 40.4 | 39.6 |
| Knowledge & Reasoning | 33.5 | 27.6 | 39.5 | 18.3 |
| Language Understanding | 60.2 | 48.6 | 35.2 | 31.1 |
| Arts & Human Taste | 28.8 | 25.8 | 37.7 | 17.8 |
kas-4b leads the 4B tier on:
- Knowledge & Reasoning (39.5) โ highest of all 4B models; higher LoRA rank (r64) likely helps on harder reasoning tasks
- Arts & Human Taste (37.7) โ highest of all 4B models by a wide margin
kas-4b trails on:
- Language Understanding (35.2) โ 13โ25 points behind ezjev and Nox; language-heavy fine-tuning data favors those models
- Tools & Automation (49.0) โ behind ezjev (69.9) and Nox (60.1)
Profile: the most balanced 4B model on the benchmark โ consistent across all five areas rather than peaking on one or two.
Training
- Base model: Qwen/Qwen3-4B (Apache 2.0, pure transformer)
- Method: LoRA fine-tune (PEFT), adapter merged into base weights
- LoRA config: rank 64, alpha 128, 2 epochs, LR 1e-4, batch size 16, 4096 token limit
- Recipe: inspired by Scion (Sinan Ozdemir)
- Hardware: NVIDIA RTX PRO 6000 Blackwell (HF Jobs)
Engine
This model is served via a custom engine (kas_engine.engine:KasEngine) that:
- Maps options to single-token uppercase labels (AโZ, then two-letter)
- Repeats the JSON prompt once (Scion-style)
- Scores label logits at the generation position, softmaxed to probabilities
- Refuses over-context prompts โ never truncates
Engine code: kashavpiya/kas-4b
Latency
Measured on RTX PRO 6000, single process, one request at a time:
| Metric | Value |
|---|---|
| Median | 39.2 ms |
| p95 | 546.4 ms |
| Mean | 153.3 ms |
Evaluation
Run with Decision Index kit v0.2.1 on 150,759 rows (184 unsupported / over-context, not truncated).
Notes
Not trained on Decision Index suite data. Base model license: Apache 2.0.
- Downloads last month
- 22