Qev-4B
Qev-4B turns context, a question and candidate answers into a decision and a probability for every option. It supports Choice, Noul (yes/no) and Score (ordered ratings) through the same API as Qev-2B and Qev-9B.
Source and installation · Training method · 中文训练说明
This release starts directly from Qwen3.5-4B-Base, with fresh LoRA and decision-head parameters. A single two-epoch stage learns a 9B teacher's option probabilities. It does not inherit an earlier distilled 4B checkpoint or include representation-response continuation.
| Property | Qev-4B v0.1.0 |
|---|---|
| Base | Qwen3.5-4B-Base, revision 710fd005d44d55ee27b7ad5147e318e546efdbfe |
| Adaptation | LoRA rank 64, alpha 128 |
| Decision head | 256 dimensions, four attention heads, two Transformer layers |
| Candidate interaction | Full sibling interaction in the final full-attention layer |
| Training | Teacher-probability cross entropy; 44,576 single-question inputs |
| Schedule | Two epochs, 2,786 steps, four GPUs, global batch 32, seed 17 |
| Temperatures | Teacher 1.563437713227029; student 1 |
| Input limits | 4,096-token state and complete path; question 512, candidate 256 |
| Computation | BF16 backbone, FP32 decision head |
| Download | Approximately 524 MiB; the Qwen base downloads separately |
Use the model
Install the Qev source package with Python 3.12 and a hardware-compatible PyTorch 2.8.0 build:
from qev import Qev
model = Qev.from_pretrained("AustinFu/Qev-4B", revision="v0.1.0", device="cuda")
answers = model.predict({
"state": "I was charged twice for the same order.",
"questions": {
"department": {
"type": "choice",
"instructions": "Which team should handle this?",
"criteria": {"billing": "Charges and refunds", "shipping": "Delivery problems"},
}
},
})
print(answers["department"]["choice"])
print(answers["department"]["probabilities"])
The loader restores the LoRA adapter, decision head, interaction gate and tokenizer, and downloads the pinned Qwen base. No teacher is needed at inference time. The package contains all trained adaptation tensors, training and fine-tuning configurations, evaluation results and license files; optimizer state and base weights are excluded.
Recorded evaluation
| Benchmark | Correct / total | Accuracy (%) |
|---|---|---|
| Decision development · clean | 1099 / 1264 | 86.95 |
| Transfer development · clean | 538 / 656 | 82.01 |
| MMLU-Pro | 503 / 1000 | 50.30 |
| SemIf · handwritten | 130 / 144 | 90.28 |
| scienthoon | 670 / 873 | 76.75 |
| WANLI | 188 / 256 | 73.44 |
| JevBench public | 190 / 231 | 82.25 |
Results use BF16 backbone computation, an FP32 head, temperature 1 and full causal reference execution. There were zero rejected questions in all seven original evaluations. The development and SemIf rows use the same subsets as the Qev leaderboard; full-suite counts are included in evaluation.json.
These are measurements from one checkpoint and seed. Training data and objectives differ across the 2B, 4B and 9B releases; their scores do not isolate model-size effects. A native 4B language-model baseline was not run. The gameplay recordings in the source repository use 9B models.
Training and availability
The input pool combines general decisions, science and reasoning, Principle judgments, controlled boundaries, web actions and additional rule/reasoning tasks. Each record contains one question with its hard label removed. Training uses a single pool without late-stage repetition or online option changes.
The recorded teacher is a separate 9B research checkpoint trained with additional web and rule tasks, at step 2,658; it is different from the public Qev-9B v0.2.0. The teacher, full 44,576-input pool and cached teacher outputs are not distributed. Qev-train remains a separate release of 2,442 synthetic examples with hard labels and synthesis documentation. The source repository provides the 4B method and a runnable example using a public teacher and user data.
Use configs/qev-4b-finetune.json with python -m qev.train --init-checkpoint AustinFu/Qev-4B@v0.1.0 for supervised adaptation. Validate performance and probabilities on your own task; calibration and production reliability have not been established.
The Qev adaptation weights use Apache-2.0. The Qwen base, tokenizer and source datasets retain their respective terms. License and attribution.
Model tree for AustinFu/Qev-4B
Base model
Qwen/Qwen3.5-4B-Base