CPM-jev

Version: v0.1-research-preview

An experimental Jev-style / System-One decision model based on MiniCPM5-2B-Base. This is an independent community project and is not affiliated with or endorsed by TypeSafe AI or OpenBMB.

CPM-jev is an independently developed LoRA fine-tune of openbmb/MiniCPM5-2B-Base, with an added decision head for scoring candidate actions. It is a research preview, not a chat model, not production-ready, and should not use generate() as its primary decision interface.

Model Description

Given a state, a question, and candidate options, CPM-jev assigns one scalar score to each option and normalizes those scores into a probability distribution:

state + question + candidate options
    -> decision scorer
    -> probability distribution

The probabilities are useful for ranking, routing, selective prediction, and confidence-aware escalation. They must not be interpreted as universally calibrated real-world probabilities.

Base model weights are not duplicated in this repository. Download openbmb/MiniCPM5-2B-Base separately.

Architecture

  • Backbone: openbmb/MiniCPM5-2B-Base
  • Adaptation: LoRA, rank 16, alpha 32, dropout 0.05
  • LoRA targets: q_proj, k_proj, v_proj, o_proj, gate_proj, up_proj, down_proj
  • Decision head: FP32 linear scalar head over the last non-padding token
  • Candidate normalization: raw softmax across the options for one question
  • Maximum sequence length: 512 tokens, left truncation

Each candidate is encoded independently with the same prompt template. The scalar scores are only comparable among the options supplied in the same call.

Training Data

CPM-jev was trained on the train split of SargeDev/jev-distill-corpus-v3, using exactly 655,806 samples. Probability targets in that dataset include data distilled from a Jev teacher.

No samples from validation, calibration, test, test_set_30k, or ood were used for training. Those splits were reserved for model selection, calibration analysis, final evaluation, and OOD evaluation as appropriate.

Training Procedure

  • Objective: soft-target cross-entropy over candidate distributions
  • Epochs: 1
  • Optimizer steps: 40,988
  • Effective batch size: 16 (micro-batch 4, gradient accumulation 4)
  • Learning rate: 1e-4
  • Weight decay: 0.01
  • Warmup ratio: 0.03
  • Seed: 42
  • Maximum length: 512
  • Trainable parameters: 25,118,721 (about 1.10% of total)

The final weights are the completed Stage 3 run. No evaluation split was mixed into training.

Evaluation

The primary final benchmark is jev-distill-corpus-v3/test_set_30k, containing 29,955 samples. Reported values use raw probabilities.

Metric Result
Accuracy 0.878317
NLL 0.727553
Brier score 0.011975
ECE 0.212610
Signed ECE -0.212610
Correct mean confidence 0.692363
Incorrect mean confidence 0.473300

Compared with the Stage 2 100k rebaseline on the same benchmark:

Metric Absolute change
Accuracy +4.934 percentage points
NLL -0.034981
Brier score -0.018019
ECE +0.034864

Accuracy improved, but calibration error worsened. This trade-off is important when deciding whether to abstain or escalate.

Selective Accuracy

Threshold Coverage Accuracy Risk
>=0.70 41.92% 99.68% 0.32%
>=0.80 27.98% 99.80% 0.20%
>=0.90 16.03% 99.90% 0.10%

These figures come from jev-distill-corpus-v3/test_set_30k. They do not imply the same accuracy on arbitrary real-world tasks, distributions, prompts, option sets, or deployment environments.

Calibration

Temperature scaling fitted on the calibration split produced T = 1.008740. It did not improve overall ECE consistently across validation, calibration, test, and test_set_30k, so the release defaults to raw probabilities (use_temperature: false).

The negative signed ECE on the primary in-domain benchmark indicates that the model is substantially under-confident there. High ECE remains a central limitation.

OOD Evaluation

On the held-out ood split (13,058 samples), raw probabilities produced:

Metric Result
Accuracy 0.878542
NLL 0.881370
Brier score 0.196659
ECE 0.093474
Signed ECE +0.093387

The positive signed ECE and much larger Brier score show a different, over-confident OOD failure mode. OOD calibration is materially weaker than the in-domain selective-accuracy table suggests.

Intended Use

Research and prototyping uses include:

  • agent tool routing
  • model routing
  • retry and recovery decisions
  • next-action selection
  • result judging
  • confidence-aware escalation
  • System-1 / System-2 routing

The model is intended to compare explicitly supplied options. Applications should define abstention and human-escalation policies and validate them on their own workload.

Limitations

  1. The model is clearly under-confident on the reported in-domain evaluation.
  2. ECE remains high.
  3. OOD calibration is materially weaker than in-domain calibration and exhibits over-confidence.
  4. OOD metrics are accuracy 0.878542, NLL 0.881370, Brier 0.196659, ECE 0.093474, and signed ECE +0.093387.
  5. The model must not be described as fully calibrated.
  6. Current evidence supports selective decision and routing use more strongly than treating its probability as an absolute real-world probability.
  7. Do not directly use its probabilities for automated medical, financial, legal, safety-critical, or other high-risk decisions.
  8. Results are specific to the released corpus and evaluation procedure; independent external evaluation is still needed.
  9. The model is not a chat model and generate() is not its decision interface.

Installation

git clone https://huggingface.co/link921/CPM-jev
cd CPM-jev
python -m venv .venv
# Windows: .venv\Scripts\activate
# Linux/macOS: source .venv/bin/activate
pip install -r requirements.txt

The first load downloads the base model unless it is already cached. You can pass a local base-model directory with base_model=... or --base-model ....

Inference Example

from inference import DecisionModel

model = DecisionModel(".")
result = model.decide(
    state="The previous tool call failed twice.",
    question="What should the agent do next?",
    options=[
        "retry",
        "switch_tool",
        "ask_user",
    ],
)

print(result)

Output schema (illustrative values):

{
  "options": [
    "retry",
    "switch_tool",
    "ask_user"
  ],
  "probabilities": [
    0.08,
    0.84,
    0.08
  ],
  "choice": "switch_tool",
  "confidence": 0.84
}

Run the included example:

python example.py

Or use the CLI:

python inference.py --model-dir . --state "The previous tool call failed twice." --question "What should the agent do next?" --options retry switch_tool ask_user

Do not call generate() to obtain the primary decision. CPM-jev compares candidate scores and applies softmax across the supplied options.

Citation

If this research preview is useful, cite the repository and the upstream model and dataset:

@software{cpm_jev_2026,
  title        = {CPM-jev: A MiniCPM5-2B Jev-style Decision Model},
  year         = {2026},
  version      = {v0.1-research-preview},
  url          = {https://huggingface.co/link921/CPM-jev}
}

License

This repository is released under the Apache License 2.0. The upstream base model and dataset are currently marked Apache-2.0 on Hugging Face. Users remain responsible for reviewing upstream licenses, dataset content, and applicable laws for their intended use.

Acknowledgements

  • Thanks to OpenBMB for MiniCPM5-2B-Base.
  • Thanks to SargeDev for jev-distill-corpus-v3.
  • The corpus contains probability targets that include data distilled from a Jev teacher. This release is an independent community project and does not claim official status or endorsement from TypeSafe AI, OpenBMB, or the dataset authors.
Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for link921/CPM-jev

Finetuned
(5)
this model

Dataset used to train link921/CPM-jev