Qev — decisions, grounded in Qwen

Qev-9B

Qev fine-tunes Qwen into a decision model. Give it context, a question, and answer options; receive a decision and a probability for each option. One model supports Choice, Noul (yes/no), and Score (ordered ratings).

Source and documentation · 中文说明 · Model weights

Qev-9B combines a general decision-training dataset with additional alignment and rule-compliance examples, repeated three times during the second half of training. It uses the full candidate-interaction architecture shown below.

Quick start

Use Python 3.12 and install a hardware-compatible PyTorch 2.8.0 build, then install Qev from its source repository:

git clone https://github.com/QiqianFu/Qev.git
cd Qev
python -m pip install -e .
from qev import Qev

model = Qev.from_pretrained(
    "AustinFu/Qev-9B", revision="main", device="cuda"
)
answers = model.predict({
    "state": "I was charged twice. Please help immediately.",
    "questions": {
        "department": {
            "type": "choice",
            "instructions": "Which team should handle this?",
            "criteria": {"billing": "Charges and refunds", "shipping": "Delivery problems"},
        }
    },
})
print(answers["department"]["choice"])
print(answers["department"]["probabilities"])

The Qev loader downloads this checkpoint and the pinned Qwen base separately. Use the Qev Python or JSONL interface to load the adapter, decision head, and interaction gate together. The default execution uses shared-prefix caching. All three task formats.

For JSONL inference:

python -m qev.predict \
  --checkpoint AustinFu/Qev-9B@main \
  --input examples/requests.jsonl --out runs/predictions.jsonl \
  --device cuda --weights-dtype checkpoint

Architecture and training

Qev encodes context, questions and answer options, then scores the options with its decision head.

Field Released model
Base Qwen/Qwen3.5-9B-Base
Adaptation LoRA rank 64, alpha 128; learned decision head and interaction gate
Decision head Shared 4096→256 projection; two 4-head Transformer layers; scalar scorer
Backbone interaction last-full-attention
Computation BF16 backbone; FP32 decision head and key reductions
Stored adaptation tensors FP32
Main training partition 34,546 records
Late partition 1,419 alignment records + 364 rule-compliance judgments
Training schedule Two epochs, global batch 32, seed 17
Late mixing Starts halfway through main training; late examples repeat three times
Selected checkpoint Step 2327

The main and late partitions intentionally share 249 replay records. Qev-train publishes 1,842 synthetic alignment, rule-compliance and world-knowledge examples with original hard labels and generation documentation. The complete mixed training corpus is not distributed. Training guide · Data recipe.

Evaluation

Qev-9B uses BF16 backbone computation; Kev-9B uses FP32.

Qev-9B and Qev-2B accuracy compared with their Qwen3.5 base models on seven benchmarks.

Benchmark Jev (reference) Qwen3.5-9B-Base Qev-9B Kev-9B
Decision development · clean 84.49 77.69 87.42 87.18
Transfer development · clean 85.67 74.39 83.99 82.16
MMLU-Pro · 1,000 83.50 50.40 54.60 51.10
SemIf · 144 handwritten 96.53 90.28 93.75 90.97
scienthoon · 873 75.26 68.84 72.28 75.49
WANLI · 256 75.78 67.97 72.66 70.31
JevBench public · 231 85.71 75.76 81.39 75.76

Accuracy (%). Bold compares Qev with Kev. JevBench is public-set accuracy: Qev answers 188 of 231 questions correctly. It is not the official JevBench composite score.

These are the recorded results for the released checkpoint, using full causal reference execution. The selected model is a single seed, and public benchmarks were observed during research iteration. The comparisons do not isolate architecture gains. Results, sources, and reproduction commands.

Files

The approximately 690 MiB inference package contains:

  • adapter/: LoRA configuration and weights.
  • head.safetensors and joint.safetensors: decision head and interaction gate.
  • model.json and tokenizer/: Qev configuration and tokenizer.
  • training_config.json and benchmarks.json: training configuration and recorded results.
  • LICENSE, NOTICE, THIRD_PARTY_NOTICES.md, and licenses/: license terms and attribution.

The adaptation tensors are byte-identical to the selected research checkpoint. Qwen base weights are fetched separately at the version recorded in the model configuration. Optimizer state is excluded; initialize new fine-tuning with --init-checkpoint AustinFu/Qev-9B.

Intended use and license

Qev supports research and development of routing, rule judgments, and rubric ratings over explicit options. Evaluate it on your application's inputs and decision thresholds. Probabilities depend on the supplied options; calibration and production reliability have not been established.

Qev's code, adaptation weights, documentation and original illustrations use Apache-2.0. Qwen models and the bundled tokenizer retain Alibaba Cloud's Apache-2.0 license. Qev adapts conventions from Jared Palmer's Kev, whose Apache-2.0 license and attribution are retained. JevBench tasks and external dependencies retain their own terms; the training corpus is not redistributed. See the included LICENSE, NOTICE, THIRD_PARTY_NOTICES.md, and licenses/ for the full texts and scope.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for AustinFu/Qev-9B

Adapter
(53)
this model

Dataset used to train AustinFu/Qev-9B