Instructions to use PelaAI/KnowLine-4B-Gen3 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use PelaAI/KnowLine-4B-Gen3 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-classification", model="PelaAI/KnowLine-4B-Gen3")# pip install -U transformers accelerate # Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("PelaAI/KnowLine-4B-Gen3") model = AutoModelForMultimodalLM.from_pretrained("PelaAI/KnowLine-4B-Gen3", device_map="auto") - Notebooks
- Google Colab
- Kaggle
KnowLine-4B-Gen3
The third release of PelaAI's 4B System One decision model, trained on a single all-in-one machine by an agent-driven data loop.
Inference guide ยท ไธญๆ่ฏดๆ ยท weights Apache-2.0 ยท previous version: KnowLine-4B-Gen2
Highlights:
Decision Index 0.3, public suite: 63.11 (self-run), 0.57 above Gen2 (62.54). For comparison: Perplexity Decider v1.1 (27B) scores 62.25 and Jev 1.13 57.96 (board), and the best โค5B model on the board has a public score of 50.82; see Comparison.
Better than Gen2 on our Jev-style and Chinese evaluations: Open-Jev 1.1 test / OOD 88.4 / 87.2 (Gen2: 87.4 / 86.6), C-Eval 79.2 (78.6). Instructions planted in the state change the answer 3.4% of the time, down from 4.0%.
Calibration: computed the board's way, the Brier score is about 0.31-0.33, roughly 3rd of 114 models on the board, and ECE is about 0.07-0.08 (board median 0.084); see Calibration.
KOF '98 harness: 16-1-1 against Jev 1.13 and 12-4-2 against StartLux-Decision-4B in single-bout mirror matches; see Game harness.
An AI agent ran the whole loop. In each round it:
- found the areas where the model was weak;
- proposed a targeted data group;
- built and decontaminated the data;
- trained and evaluated on it;
- kept or rejected the change based on the evidence.
For this round, see What changed in Gen3.
KnowLine is an independent model. It is not affiliated with, endorsed by, or derived from TypeSafe or Jev.
What it does
You send a state (text or a chat) and up to 64 typed questions: yes/no, choose one of k, or score on a rubric. The model answers every question with a probability distribution over its options:
- one forward pass per question, with no generated text;
- the output is the probability of each option's label token;
- existing Jev clients only need a new base URL; Gen1 and Gen2 users only change the model name, since the interface and serving settings are the same.
| Base model | Qwen/Qwen3.5-4B (Apache-2.0) |
| Training | LoRA SFT (rank 32, alpha 64, language model only), merged into bf16 weights. |
| Release format | bf16 weights. The vision tower and MTP head are the base model's; config, tokenizer and chat template are identical to Gen1 and Gen2. |
| Languages | English, Simplified Chinese, Traditional Chinese |
Quickstart
pip install "sglang==0.5.21" "transformers==5.12.1" requests
bash serve_knowline.sh PelaAI/KnowLine-4B-Gen3 0 8080 # SGLang (FP8 at load) on :9080 + /v1/systemone on :8080
curl -s http://127.0.0.1:8080/v1/systemone -H 'Content-Type: application/json' -d '{
"model": "m",
"state": "Customer: my order arrived broken, I want my money back.",
"questions": {"refund": {"type": "noul", "instructions": "Should the agent offer a refund?"},
"tone": {"type": "choice", "instructions": "Customer tone?",
"criteria": {"angry": "Angry", "neutral": "Neutral", "happy": "Happy"}}}}'
- Without SGLang:
python knowline_server.py --model PelaAI/KnowLine-4B-Gen3 --backend hf --port 8080uses transformers only. It is slower and runs bf16. - Front end:
knowline_server.pyis a single file and needs only transformers and requests. It is the same file as in Gen2. - Full settings: the exact settings of our Decision Index run are in INFERENCE.md.
What changed in Gen3
- Finding the gaps. Gen2 was still weakest on knowledge and reasoning (42.7).
- New data. On top of the previous round's data, a knowledge replay group: multiple-choice items from public train splits covering general knowledge, science, medicine, math word problems and logic, in English and Chinese. None of them comes from a Decision Index benchmark, and the new rows went through decontamination again.
- Training. One epoch, with the learning rate decayed all the way down. The replay raised knowledge and reasoning a little but cost BPoMP in the arts area.
- Selection. This round compared 13 candidates, which scored 62.54-63.25 on the 0.3 public suite. Gen3 scores 63.11 on that suite, within noise of the highest score. Of the 13 candidates it was the best on our Jev-style evaluations (JevBench 86.6, JevBench-hard 72.1) and C-Eval (79.2), and near the top on out-of-distribution generalisation (Open-Jev 1.1 OOD 87.2).
- Where the gain comes from. The gain over Gen2 is spread over many benchmarks, led by NLI4CT (57.6 โ 62.3), CRUXEval (38.8 โ 44.0), WinoGrande (73.8 โ 78.8), RAGTruth (53.3 โ 56.8) and GSM8K (58.0 โ 61.5). The train splits of WinoGrande and GSM8K are in the training data in Decision Index request format. BPoMP (59.4 โ 54.5) and PhishNChips (95.0 โ 92.0) dropped.
Comparison
Decision Index 0.3, public suite. The full 0.3 score adds private tests that only the maintainers run (0.5 same-skill, 0.3 new-domain); ours is not available yet.
| model | size | DI 0.3 public | DI 0.3 full | source |
|---|---|---|---|---|
| KnowLine-4B-Gen3 | 4B | 63.11 | not yet scored | self-run, official kit |
| KnowLine-4B-Gen2 | 4B | 62.54 | not yet scored | self-run, official kit |
| KnowLine-4B-Gen1 | 4B | 60.47 | not yet scored | self-run, official kit |
| Perplexity Decider v1.1 | 27B | 62.25 | 62.75 | board |
| Clef | 27B | 61.71 | 53.08 | board |
| Jev 1.13 | (API) | 57.96 | 60.11 | board |
| RSI-Jev v6.1-VL | 4B | 50.98 | not listed | self-reported |
| ezjev 4B s2 | 4B | 50.82 | 46.95 | board |
| jiwo 4B | 4B | 45.76 | 42.86 | board |
| Nox 4B | 4B | 44.21 | 44.95 | board |
Public-suite scores by area (all 0.3 public):
| area | KnowLine-4B-Gen3 | KnowLine-4B-Gen2 | Perplexity Decider v1.1 (27B) | Jev 1.13 |
|---|---|---|---|---|
| Knowledge & reasoning | 43.9 | 42.7 | 52.0 | 53.9 |
| Language | 64.8 | 63.3 | 67.5 | 59.2 |
| Retrieval & routing | 70.9 | 71.3 | 61.3 | 55.4 |
| Tools & agents | 86.6 | 86.6 | 78.9 | 75.1 |
| Arts & taste | 49.8 | 50.2 | 47.0 | 39.1 |
Releases
The model is trained in a self-evolving loop, and new versions will follow.
| model | date | DI 0.3 public | DI 0.2.1 | golden held-out (en / zh-Hans / zh-Hant) | notes |
|---|---|---|---|---|---|
| KnowLine-4B-Gen3 (this model) | 2026-10-09 | 63.11 | โ | 69.4 / 70.8 / 65.0 | third release |
| KnowLine-4B-Gen2 | 2026-10-08 | 62.54 | โ | 69.7 / 71.2 / 64.8 | second release |
| KnowLine-4B-Gen1 | 2026-10-07 | 60.47 | 60.92 | 68.7 / 69.9 / 64.1 | first release |
From Gen2 on we run Decision Index 0.3 only, not 0.2.1.
Evaluation
All results are self-run and not verified by a third party.
Held-out and Chinese evaluations
| suite | Gen3 | Gen2 |
|---|---|---|
| In-house evaluation set, English | 69.4 | 69.7 |
| In-house evaluation set, Simplified Chinese | 70.8 | 71.2 |
| In-house evaluation set, Traditional Chinese | 65.0 | 64.8 |
| C-Eval (4 categories, macro) | 79.2 | 78.6 |
| Open-Jev 1.1 test / OOD | 88.4 / 87.2 | 87.4 / 86.6 |
| Prompt injection: answers changed (lower is better) | 3.4% | 4.0% |
Web operation (Mind2Web official test splits, evaluation only)
| split | Gen3 element selection | Gen3 operation (balanced) | Gen2 element selection | Gen2 operation (balanced) |
|---|---|---|---|---|
| test_task (websites seen in training, new tasks) | 92.7 | 97.0 | 92.9 | 97.0 |
| test_website (new websites) | 91.3 | 97.3 | 90.8 | 97.3 |
| test_domain (new domains) | 91.8 | 97.2 | 91.7 | 97.4 |
- About the same as Gen2; the differences are within noise.
- So far this is only used to explore the model in RPA-style automation and to check that it generalises to some degree.
- The task is to pick the target element among it and up to 5 other candidates sampled from the page. This is easier than the original Mind2Web protocol, so do not compare it with the Mind2Web leaderboard.
Game harness (KOF '98)
- Setup: single-bout character-mirror matches, 18 games per pair, argmax actions; the same settings as Gen1's round robin and Gen2's matches.
- Result: 16-1-1 against Jev 1.13, 12-4-2 against StartLux-Decision-4B and 4-14 against Gen2; 32-19-3 overall, score 0.620 [0.49, 0.74].
- Play style: it picks moves by distance: mostly the 623C anti-air uppercut up close (75%), special_2 and heavy kick at mid range, and almost only special_2 from far away. Overall it uses the uppercut 23% of the time.
- Caveats:
- 18 games per pair give wide intervals.
- 37 of the 54 games ended at time-out, so most wins are on remaining health. Gen3 won 8 games by K.O. (2 against Jev, 4 against StartLux, 2 against Gen2).
- This is a measured result in this harness, not general fighting-game skill.
Calibration
Computed on the 0.3 public suite with the method the Decision Index board describes: each field is right or wrong, confidence is the probability on the chosen option, benchmarks are weighted equally, and there are 10 equal-width bins.
| metric | Gen3 | Gen2 | board median (114 models) | Jev 1.13 |
|---|---|---|---|---|
| ECE (lower is better) | 0.07-0.08 | 0.06-0.07 | 0.084 | 0.074 |
| Brier score (lower is better) | 0.31-0.33 | 0.31-0.33 | 0.49 | 0.36 |
| Confidence โฅ95% but wrong | 2.7% | 2.2% | 1.3% | 2.1% |
| Mean confidence / accuracy | 0.84 / 0.76 | 0.83 / 0.76 | 0.81 / 0.74 |
- Why ranges: the board does not say which 32 benchmarks it uses. We checked our computation against three board models whose results are public. Our ECE was within about 0.015 of the board's, so we give ranges over the plausible benchmark sets.
- Still overconfident overall: mean confidence is about 7-8 points above accuracy. Most of the gap is on hard reasoning (HLE, CRUXEval, MuSR) and humour (Humicroedit, New Yorker caption matching).
- Recommendation: if you act on probability thresholds, fit a temperature on your own data.
Disclosures
- Decision Index format training data: Gen3 comes from three training rounds. In each, about 25-31% of the data is in Decision Index request format. This includes train splits of public datasets rewritten in that format (for example GSM8K, WinoGrande XL and ACOS), and synthetic items written in the same format. No Decision Index test item is included.
- Decontamination:
- Every component of a training row of at least 60 characters (a line or paragraph) was checked against the components of every Decision Index row and all of our evaluation sets.
- Rows with a component identical to an evaluation component, or with the same first 50 characters, were removed. Text that appears in more than 4 evaluation items counts as boilerplate and is not matched.
- None of the rows new in this round was flagged; the rest of the data had already been decontaminated in earlier rounds.
- Selection on evaluations: of the 13 averaging variants, we picked one within noise of the highest Decision Index 0.3 public score that was also best on our Jev-style evaluations.
- Game data: labels come from simulator rollouts or engines. Seeds and start states are disjoint from the evaluation states.
- No model outputs as labels: no output of Jev or any other decision model was used as a training label.
- Teacher-labelled synthetic data: LLM teachers wrote and labelled our synthetic tasks. The labels were filtered, but not all were checked by a human.
Limitations
- Knowledge-heavy reasoning: weaker than larger models. The knowledge area is 43.9 (0.3); MMLU-Pro moved little (51.7) and HLE stays at chance level.
- Math: answered without reasoning. The rebuilt 0.3 GSM8K scores 61.5; the gain comes from training on GSM8K train rewritten in the same format.
- Prompt injection: an instruction planted in the state changes the answer about 3% of the time on our injection set. Keep untrusted text clearly delimited.
- Calibration: overconfident overall; see Calibration.
- Private tests: part of our public-suite advantage comes from adapting to the question formats. Some of Gen3's largest gains over Gen2 are on benchmarks whose train splits were trained on, so the 0.3 private tests may score lower.
Citation
@misc{knowline4bgen3,
title = {KnowLine-4B-Gen3: a 4B decision model},
author = {PelaAI},
year = {2026},
url = {https://huggingface.co/PelaAI/KnowLine-4B-Gen3}
}
License
- Weights: Apache-2.0, the same as the base model.
- Code:
knowline_server.pyis MIT (see the file header).
- Downloads last month
- 5