Instructions to use CVLTAI/s1-jev-cvltist with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use CVLTAI/s1-jev-cvltist with PEFT:
from peft import PeftModel from transformers import AutoModelForCausalLM base_model = AutoModelForCausalLM.from_pretrained("Qwen/Qwen3.5-4B") model = PeftModel.from_pretrained(base_model, "CVLTAI/s1-jev-cvltist") - Notebooks
- Google Colab
- Kaggle
s1-jev-cvltist β Model Card
A local, open, Jev-compatible "System One" decision model.
s1-jev-cvltist is the name of this release model (project codename "system one", s1;
the internal checkpoint id is s1-v6).
Base: Qwen3.5-4B (Apache-2.0) Β· Adapter: LoRA r=16, Ξ±=32 Β· Context: 4096
Interface: the TypeSafe Jev wire spec (choice / score / noul) Β· Size: 124 MB adapter
Release checkpoint: checkpoints/s1-v6 Β· Eval holdout: 35,594 frozen rows
Canonical home:
CVLTAI/s1-jev-cvltistβ this repo is a personal mirror of that release. Load:PeftModel.from_pretrained(base, "AEECollier/s1-jev-cvltist")overQwen/Qwen3.5-4B. The full local pipeline (training, eval harness, calibration, serving) lives at github.com/cvlt-ai/decision-model; the adapter here is self-contained (124 MB LoRA + tokenizer).
1. What this is
system one (s1) is a small decision model trained to answer Jev-style questions
(choice, score, noul) over a free-form "state", returning calibrated
probabilities rather than a bare answer. It is deliberately built on a general
instruction base (Qwen3.5-4B) with a thin LoRA adapter, so it carries knowledge
a purpose-built encoder can't, while the readout (single-token logit scoring over
option keys) keeps inference cheap and the probabilities well-formed.
The design goal was not just to match TypeSafe Jev's accuracy, but to be open and auditable, and to be structurally strong where Jev and its open clones are weak: calibration, negation/invariance consistency, and coverage (knowing when the evidence is insufficient).
2. Task
Given a state (free-form text / structured context) and one or more questions,
return a well-formed answer per the Jev wire spec:
| type | answer |
|---|---|
choice |
the chosen option key + full probability vector over options |
score |
a score on a 2β10 level scale + probability distribution |
noul |
a probability in [0,1] (the "yes/no/unless" branch) |
Up to 255 options per choice (vendor spec). Answers carry probabilities, not just the top pick β that is what makes calibration and coverage meaningful downstream.
3. Model
- Base:
Qwen/Qwen3.5-4B(Apache-2.0, 262k native context). - Adapter: LoRA (
r=16,alpha=32, dropout 0.05) on every attention projection (q/k/v/o) and MLP projection (gate/up/down) plus the gatedin_projmodules. ~124 MB. Inference-mode, bf16. - Readout:
LogitReadout(eval/readout.py) β teacher-forces each option key and reads the single-token logit; option keys are scored in chunks (key_batch) to bound peak memory. This is the same NLL proper scoring rule used at training time.
4. Training data
mixture_v6 β 264,162 choice/score/noul rows, 1 epoch, 8,077 optimizer steps
(effective batch 32, 4096 context), loss β ~0.06β0.28.
| family | rows | % | source |
|---|---|---|---|
| tsi | 128,000 | 48.5% | TaskSource-instruct (129 tasks, NLI/counterfactual/knowledge, tsi-perm commercial-safe subset) |
| go_emotions | 86,820 | 32.9% | GoEmotions (multi-label, order-augmented) |
| banking77 | 19,986 | 7.6% | Banking77 (77-way intent, order-augmented, key-only render) |
| boolq | 18,854 | 7.1% | BoolQ (CC-BY-SA; negation pairs built on it) |
| severity | 4,946 | 1.9% | bespoke severity scale |
| nimble | 4,464 | 1.7% | Bespoke Nimble c2d contrastive pairs (evidence-flip) |
| injection | 1,092 | 0.4% | prompt-injection robustness |
License posture (release line): permissive / commercial-safe. Qwen3.5-4B is
Apache-2.0; TSI is the commercial-safe subset; BoolQ is CC-BY-SA and enters only via
the --include-sa gate. Non-commercial rows (anli, multi_nli, snli, ai2_arc,
hellaswag) are quarantined out of the release mixture.
5. Evaluation β release numbers (full 35,594-row frozen holdout)
Frozen holdout, 9 families, s1-jev-cvltist at 1024 context, T=1.0. This is the citable number (the 300/family sample used during iteration is superseded by the full run).
| metric | s1-jev-cvltist (release) |
|---|---|
| macro accuracy (unweighted, 9 families) | 0.7747 |
| mmlu_pro (n=12,032) | 0.438 |
| negation violation (lower = more consistent) | 0.022 |
| ECE @ raw T=1 (pooled) | 0.105 |
| ECE @ best temperature (T=1.5) | 0.0169 |
| Brier @ best temperature | 0.1441 |
Per-family accuracy (full holdout)
| family | n | accuracy |
|---|---|---|
| injection | 126 | 0.9921 |
| banking77 | 3,076 | 0.9324 |
| boolq | 3,270 | 0.9128 |
| negation | 3,270 | 0.9110 |
| pubhealth | 7,929 | 0.7874 |
| baserate | 27 | 0.7407 |
| severity | 437 | 0.6568 |
| go_emotions | 5,427 | 0.6005 |
| mmlu_pro | 12,032 | 0.4383 |
Strong on six families (β₯0.65): injection, banking77, boolq, negation, pubhealth, baserate. The two gaps β go_emotions (0.601, the weakest large family) and mmlu_pro (0.438, the knowledge gap) β are the honest shortfalls.
Calibration
The raw pooled ECE (0.105) is high only because mmlu_pro (12k of 35.5k rows, where the model is genuinely uncertain) dominates the pool β every individual family's raw ECE is fine. Offline temperature scaling (the standard post-hoc phase) at T=1.5 brings pooled ECE to 0.0169 β the best calibration held across any version, well under the 0.03 gate β with Brier 0.1441.
JevBench (public-231, harness's own score_task)
| model | all-public | hard tier |
|---|---|---|
| s1-jev-cvltist @4096 | 0.7749 | 62/111 |
| s1-v5 | 0.7489 | 59/111 |
| AlexWortega/openjev v5 | 0.814 | 69/111 |
| Jev 1.13 (hosted) | 0.866 | 81/111 |
s1-jev-cvltist (coral) is the release checkpoint; s1-v3βv6 is our local progression. openjev (0.814) and Jev 1.13 (0.866) are external, documented numbers, not local runs. Gap to openjev closed from 0.065 to 0.039; Jev 1.13 remains 0.091 ahead.
Evidence-sensitivity (VitaminC flip probe, 200 conflict families)
| s1-jev-cvltist | |
|---|---|
| overall | 0.799 |
| SUPPORTS | 0.919 |
| REFUTES | 0.713 |
| NOT ENOUGH INFO | 0.582 |
| lazy_rate (lower = better) | 0.090 |
The broad NLI/counterfactual coverage in the TSI data taught the model to return "not enough info" when the evidence genuinely runs out, instead of forcing an answer β the "don't over-claim" property that is the whole point of a coverage-aware decision model.
6. How it got here (version trajectory, frozen 300/family suite)
| version | macro | negation | Brier | mmlu_pro |
|---|---|---|---|---|
| Laya (encoder baseline) | 0.410 | 0.710 | 0.336 | 0.137 |
| s1-v1 | 0.592 | 0.032 | 0.140 | 0.303 |
| s1-v2 | 0.710 | 0.268* | 0.161 | 0.320 |
| s1-v3 | 0.727 | 0.030 | 0.146 | 0.300 |
| s1-v4 | 0.723 | 0.022 | 0.157 | 0.253 |
| s1-v5 | 0.750 | 0.021 | 0.125 | 0.350 |
| s1-jev-cvltist | 0.779 | 0.018 | 0.123 | 0.487 |
* v2's negation spike was a training-data bug (built without BoolQ); fixed in v3.
Each line is a deliberate data decision, documented in research/05βresearch/11.
The v6 jump is TSI breadth: it closed most of the knowledge gap and improved the
differentiator axes (negation, coverage) simultaneously.
Laya (the purpose-built encoder we set out to beat) is 0.410; s1-jev-cvltist is 0.775 β +0.365 over the encoder baseline, and the largest single jump is the v5βv6 TSI-breadth step.
7. How to run it
Minimal (logit readout, no server):
from eval.readout import LogitReadout
ro = LogitReadout(ckpt="checkpoints/s1-v6", base="Qwen/Qwen3.5-4B",
device="cuda:0", temperature=1.5, max_len=1024, key_batch=8)
answers = ro.call(state_text, {"q1": {"type": "choice", "options": {...}}})
# answers["q1"] = {"choice": key, "probabilities": {key: p, ...}}
HTTP (Jev wire spec, src/s1/schema.py speaks the request/response format):
python scripts/serve_s1.py --ckpt checkpoints/s1-v6 --base Qwen/Qwen3.5-4B \
--device cuda:0 --temperature 1.5 --port 8000
# POST /v1/answer { "state": "...", "questions": { ... } }
Serving note:
LogitReadout+schema.pyare the working pieces;serve_s1.pywraps them in FastAPI. There is no batching/continuous-batching layer yet β that is the next serving step if throughput matters.
8. Limitations (honest)
- Knowledge is the gap. mmlu_pro 0.438 vs Jev's documented 0.83. This is the one axis we are clearly behind, and it is bounded by the TSI per-task extraction cap (500/task) which throttled the knowledge-heavy tasks.
- go_emotions 0.601 is the weakest large family β multi-label emotion is hard and it is only in the mix via order-augmentation.
- baserate (n=27) and injection (n=126) are small families; their per-family numbers carry wide confidence intervals.
- Raw pooled ECE is dominated by the uncertain knowledge family; always report the temperature-scaled ECE (0.0169) or per-family ECE, not the raw pool.
- Numbers are from a frozen 9-family holdout + JevBench public-231. They are not measured on Jev's private/internal suites.
9. Reproducibility
- Release checkpoint:
checkpoints/s1-v6(commit554debd..f9f229e). - Release eval:
results/s1v6_release_full.json(35,594 rows) +results/calibration_s1v6_release.json. - Build:
scripts/train_v6.sh(training),scripts/eval_v6.sh(eval chain),scripts/tsi_extract.py(TSI data). - Research log:
research/01βresearch/11(one doc per milestone). - Frozen holdout:
data/holdout/*.parquet+MANIFEST.json(sha256-pinned).
10. License
Base model Apache-2.0; release mixture permissive/commercial-safe (see Β§4). Adapter
weights released under the project license (TBD β see LICENSE).
- Downloads last month
- 31