Qev-9B
Qev fine-tunes Qwen into a decision model. Give it context, a question, and answer options; receive a decision and a probability for each option. One model supports Choice, Noul (yes/no), and Score (ordered ratings).
Source and documentation · 中文说明 · Model weights
Qev-9B combines a general decision-training dataset with additional alignment and rule-compliance examples, repeated three times during the second half of training. It uses the full candidate-interaction architecture shown below.
Quick start
Use Python 3.12 and install a hardware-compatible PyTorch 2.8.0 build, then install Qev from its source repository:
git clone https://github.com/QiqianFu/Qev.git
cd Qev
python -m pip install -e .
from qev import Qev
model = Qev.from_pretrained(
"AustinFu/Qev-9B", revision="main", device="cuda"
)
answers = model.predict({
"state": "I was charged twice. Please help immediately.",
"questions": {
"department": {
"type": "choice",
"instructions": "Which team should handle this?",
"criteria": {"billing": "Charges and refunds", "shipping": "Delivery problems"},
}
},
})
print(answers["department"]["choice"])
print(answers["department"]["probabilities"])
The Qev loader downloads this checkpoint and the pinned Qwen base separately. Use the Qev Python or JSONL interface to load the adapter, decision head, and interaction gate together. The default execution uses shared-prefix caching. All three task formats.
For JSONL inference:
python -m qev.predict \
--checkpoint AustinFu/Qev-9B@main \
--input examples/requests.jsonl --out runs/predictions.jsonl \
--device cuda --weights-dtype checkpoint
Architecture and training
| Field | Released model |
|---|---|
| Base | Qwen/Qwen3.5-9B-Base |
| Adaptation | LoRA rank 64, alpha 128; learned decision head and interaction gate |
| Decision head | Shared 4096→256 projection; two 4-head Transformer layers; scalar scorer |
| Backbone interaction | last-full-attention |
| Computation | BF16 backbone; FP32 decision head and key reductions |
| Stored adaptation tensors | FP32 |
| Main training partition | 34,546 records |
| Late partition | 1,419 alignment records + 364 rule-compliance judgments |
| Training schedule | Two epochs, global batch 32, seed 17 |
| Late mixing | Starts halfway through main training; late examples repeat three times |
| Selected checkpoint | Step 2327 |
The main and late partitions intentionally share 249 replay records. Qev-train publishes 1,842 synthetic alignment, rule-compliance and world-knowledge examples with original hard labels and generation documentation. The complete mixed training corpus is not distributed. Training guide · Data recipe.
Evaluation
Qev-9B uses BF16 backbone computation; Kev-9B uses FP32.
| Benchmark | Jev (reference) | Qwen3.5-9B-Base | Qev-9B | Kev-9B |
|---|---|---|---|---|
| Decision development · clean | 84.49 | 77.69 | 87.42 | 87.18 |
| Transfer development · clean | 85.67 | 74.39 | 83.99 | 82.16 |
| MMLU-Pro · 1,000 | 83.50 | 50.40 | 54.60 | 51.10 |
| SemIf · 144 handwritten | 96.53 | 90.28 | 93.75 | 90.97 |
| scienthoon · 873 | 75.26 | 68.84 | 72.28 | 75.49 |
| WANLI · 256 | 75.78 | 67.97 | 72.66 | 70.31 |
| JevBench public · 231 | 85.71 | 75.76 | 81.39 | 75.76 |
Accuracy (%). Bold compares Qev with Kev. JevBench is public-set accuracy: Qev answers 188 of 231 questions correctly. It is not the official JevBench composite score.
These are the recorded results for the released checkpoint, using full causal reference execution. The selected model is a single seed, and public benchmarks were observed during research iteration. The comparisons do not isolate architecture gains. Results, sources, and reproduction commands.
Files
The approximately 690 MiB inference package contains:
adapter/: LoRA configuration and weights.head.safetensorsandjoint.safetensors: decision head and interaction gate.model.jsonandtokenizer/: Qev configuration and tokenizer.training_config.jsonandbenchmarks.json: training configuration and recorded results.LICENSE,NOTICE,THIRD_PARTY_NOTICES.md, andlicenses/: license terms and attribution.
The adaptation tensors are byte-identical to the selected research checkpoint. Qwen base weights are fetched separately at the version recorded in the model configuration. Optimizer state is excluded; initialize new fine-tuning with --init-checkpoint AustinFu/Qev-9B.
Intended use and license
Qev supports research and development of routing, rule judgments, and rubric ratings over explicit options. Evaluate it on your application's inputs and decision thresholds. Probabilities depend on the supplied options; calibration and production reliability have not been established.
Qev's code, adaptation weights, documentation and original illustrations use Apache-2.0. Qwen models and the bundled tokenizer retain Alibaba Cloud's Apache-2.0 license. Qev adapts conventions from Jared Palmer's Kev, whose Apache-2.0 license and attribution are retained. JevBench tasks and external dependencies retain their own terms; the training corpus is not redistributed. See the included LICENSE, NOTICE, THIRD_PARTY_NOTICES.md, and licenses/ for the full texts and scope.
Model tree for AustinFu/Qev-9B
Base model
Qwen/Qwen3.5-9B-Base