EL JEVO 0.8B — V17

EL JEVO is a small AI model for judging answers and making decisions from text. You give it a question, the relevant evidence or rules, and a set of possible answers. It evaluates how well each answer is supported and returns a structured decision with a score for each candidate. It runs locally, including on CPU.

EL JEVO 0.8B is a further fine-tune of OpenJev by AlexWortega, built on Qwen3.5-0.8B. The project experiments with additional training for applying written rules, interpreting facts and comparing answers while retaining the original model's ability to judge whether a claim follows from a passage.

Task What you supply What it returns
Yes/no (noul) Evidence and a question with yes/no answers A selected answer and scores for yes and no
Multiple choice (choice) Evidence, a question and described answer options A selected option and scores for the candidates
Numeric scale (score) Evidence, a written rubric and allowed numeric values A selected value, candidate scores and an expected score
Natural language inference (NLI) A passage and a claim about it Entailment (the claim follows), contradiction (it conflicts), or neutral (the passage leaves it undetermined)

For example, an application could ask EL JEVO to check a proposed action against a policy. Here is an illustrative task, with an expected answer that follows from the stated rule:

Evidence: A build may be released only when every test passes.
          This build has 12 passing tests and 1 failing test.
Question: Can this build be released under the policy?
Possible answers: yes / no
Expected answer: no

The application supplies the evidence and answer choices. EL JEVO can serve as a local evaluator for candidate answers or proposed workflow actions; its scores express the model's assessment and are not guaranteed calibrated probabilities of correctness.

V17 is the current recommended EL JEVO release and the default checkpoint on main. Use revision='V17', subfolder 0.8b, to pin this release.

The complete merged BF16 checkpoint has 852,988,992 parameters (nominal 0.8B, including inherited unused vision weights), approximately 1.7 GB of stored weights, and runs on CPU. Text inference is evaluated; a GPU is optional. Detailed V17 checkpoint card.

Versions

Revision Role
V17 Current recommended release; default weights on main
V13 Historical predecessor and V17 training parent
V12 Historical broader-retention alternative; some older scores differ
V10 Historical older-rule retention alternative
V9 Original published reference

These are version tags of EL JEVO 0.8B, each stored under 0.8b/. V17 continues V13. The older cards and comparisons document earlier experiments; V13 is no longer the current recommended release. V17's recommendation reflects its development-selected checkpoint, modest gains on the sampled structural curriculum, and measured retention. It does not mean V17 wins every suite or generally beats every other model.

Evaluation

V17 compared with its V13 predecessor

The fixed V13, V16 and V17 checkpoints ran on 35,580 identical cases across 22 suites, on one RTX PRO 4500 Blackwell with the same BF16, context and batching settings. Development selected V17 at update 1,000 before the final confirmation. The new confirmation contains 445 roots, 1,335 total views and 86 held-out structural compositions.

V17 promotion confirmation V13 V16 experiment V17
Primary decisions (890) 62.13% 62.58% 63.82%
Both decisive answers correct (445 roots) 33.93% 34.83% 37.53%
Invariant views (445) 62.70% 61.80% 62.70%
Public JevBench (231) 71.00% 70.56% 71.00%

V17's primary gain over V13 is +1.69 percentage points: exploratory paired 95% interval [+0.22, +3.15] by root and [+0.32, +3.24] by structural composition. Both-correct decisive pairs improve by 3.60 points. Public JevBench remains 164/231 (71.00%). All 21 older suites stay within the preset two-percentage-point accuracy or 3% relative numeric-MAE regression limits; passing those limits allows small regressions. One training seed and shared language themes limit broader conclusions.

Mechanism Primary decisions V13 V17
bounded uncertainty 230 61.74% 63.04%
causal interventions 240 48.33% 48.75%
joint allocation 240 67.92% 68.75%
quantified collections 180 73.33% 78.33%

Causal intervention remains difficult, at 48.75% in this confirmation. Ordered scoring is measured in older suites; the four new mechanisms use binary and choice questions.

Older regression suite Cases V13 V17
Public NLI sample 396 64.39% 64.39%
MNLI holdout 120 88.33% 88.33%
Operational probes 324 87.04% 87.04%
LLMBar 838 52.98% 53.34%
JudgeBench 1240 59.03% 59.03%
ContractNLI 2091 61.36% 61.07%
HelpSteer2 argmax numeric MAE ↓ 2076 1.076 1.075

All 22 suites and paired intervals · Machine-readable promotion results · Development selection · Training experiment

V17 compared with original OpenJev 0.8B

A later completed inference-only follow-up, executed during the V18 round, compares the same published V17 weights with the original upstream OpenJev on the older V13 structural holdout. Both used the same GPU, BF16, context and batching. This is a separate 640-decision comparison; its scores are not mixed with the 890-decision confirmation above.

Matched structural follow-up Original OpenJev 0.8B V17
Primary decisions (640) 281/640 — 43.91% 405/640 — 63.28%
Both decisive answers correct (320 roots) 13/320 — 4.06% 118/320 — 36.88%
Invariant views (320) 140/320 — 43.75% 206/320 — 64.38%

The primary gain is +19.38 percentage points, with a paired root-bootstrap 95% interval [+15.31, +23.28]. The four mechanism families and language themes are related to EL JEVO training. This older observed structural test is regression evidence; the result is a targeted task gain, not independent source transfer or universal superiority. Other external peers have not been evaluated in this follow-up. Complete comparison and boundaries · Checkpoint pins and aggregate evidence.

Historical comparisons with other small models

The full earlier comparison remains available: 20 checkpoints on 24,270 identical inputs, including OpenJev, Laya, Julia, Open-JEV DeBERTa, Lumma, Kev, Tiny-Jev, FineCat, EttinX, ModernCE, GLiNER2.5-Decide and EL JEVO V9–V13. V17 was not evaluated in that 20-checkpoint campaign. Its separate comparisons are reported above.

Show the earlier 20-checkpoint table and its evaluation limits

Historical 20-checkpoint comparison (V17 is not in this table)

20 fixed checkpoints, 24,270 identical cases across 16 suites. The upstream OpenJev 0.8B is the primary baseline. The table also includes Laya (three variants), Julia, Open-JEV DeBERTa, Lumma (two sizes), Kev (two sizes), Tiny-Jev, FineCat, EttinX, ModernCE (two sizes), GLiNER2.5-Decide and EL JEVO V9/V10/V12/V13. Every prediction file and input subset was verified before scoring.

Model JevBench (231) NLI (396) V10 rules (2,880) LLMBar (838) JudgeBench (1,240) ContractNLI (2,091) Prediction coverage
EL JEVO V13 (historical) 71.00% 64.39% 41.56% 52.98% 59.03% 61.36% 100.00%
OpenJev 0.8B 71.43% 66.92% 34.58% 49.05% 56.61% 60.98% 100.00%
EL JEVO V12 71.86% 65.15% 41.22% 55.61% 59.76% 61.65% 100.00%
EL JEVO V9 71.43% 65.15% 39.97% 49.76% 57.26% 61.93% 100.00%
EL JEVO V10 70.56% 64.39% 37.88% 50.36% 56.85% 62.41% 100.00%
laya 48.05% 63.13% 33.26% 40.10% 0.08% 1.24% 81.11%
laya-typed-decisions 50.22% 60.61% 32.43% 49.76% 20.81% 5.93% 88.54%
laya-multilingual 45.45% 55.81% 33.61% 50.72% 51.61% 44.57% 99.50%
julia-1 44.59% 32.58% 33.33% 47.61% 49.11% 44.43% 99.50%
open-jev-deberta-v3-large 48.48% 33.08% 33.78% 27.80% 0.00% 0.00% 76.63%
lumma-fev-0.6b 46.75% 35.86% 32.64% 49.28% 49.60% 37.35% 99.07%
lumma-fev-0.1b 37.66% 33.84% 34.65% 49.28% 34.84% 13.15% 92.17%
kev-0.8b 63.20% 44.19% 35.73% 54.42% 54.76% 48.59% 99.09%
kev-0.5b 48.92% 40.66% 33.16% 52.03% 49.60% 38.02% 99.16%
tiny-jev 60.61% 45.20% 34.86% 56.92% 53.79% 41.22% 98.95%
finecat-nli-xxs 47.19% 57.07% 33.33% 53.46% 50.40% 20.76% 99.30%
ettinx-nli-xxs 47.62% 43.43% 33.12% 58.11% 52.42% 23.72% 99.30%
modernce-base-nli 46.75% 56.82% 33.47% 49.52% 44.52% 40.03% 95.88%
modernce-large-nli 50.65% 65.40% 33.78% 42.72% 46.45% 48.78% 95.88%
GLiNER2.5-Decide (512) 57.14% 34.34% 34.06% 40.10% 49.35% 44.62% 100.00%

On these saved runs, V13 scores 164/231 on public JevBench versus OpenJev's 165/231, and 64.39% on the NLI sample versus 66.92%. V13's V10-rule score is 41.56% versus 34.58%. This is a targeted rule-task gain; it does not establish general superiority over OpenJev or every peer.

Accuracy uses all cases; unsupported inputs count as wrong. The report also provides the same accepted subset for every model. Prediction coverage does not imply full evidence coverage: GLiNER uses a 512-token cap, and native peers have different input limits. These are saved quality measurements on identical inputs from separate inference runs (RTX 4090, RTX PRO 4500 Blackwell and A40), not one hardware-controlled run or a latency ranking. The V10 rules are older regression tests with generator families related to EL JEVO training.

All suites, common support and answer types · JSON · CSV · Original 17-model report

At the time of this historical table, the V13 structural report included GLiNER2.5-Decide but lacked original OpenJev. The later V17 versus OpenJev follow-up completes that upstream comparison using published V17 on the same structural inputs. The other external peers still lack completed structural-holdout results.

Show historical V12/V13 structural comparisons

Historical V12/V13 structural comparisons

Same matched run: new structural holdout V9 V13
Primary decisions (640) 44.69% 62.81%
Invariant views (320) 44.06% 61.88%
Both decisive answers correct (320 pairs) 3.44% 35.31%

The paired primary gain is 18.125 percentage points, with a 95% root-cluster interval 14.375–22.031 points. Both fixed models used identical BF16/context/batching on the same RTX PRO 4500 Blackwell worker. The cases contain 320 roots and 83 held-out structural compositions in four mechanisms; language themes are shared across splits. This supports a targeted curriculum gain, not general decision-making superiority. Direct comparison.

Original 20-suite comparison V12 V13 GLiNER2.5-Decide
JevBench public (231) 71.86% 71.00% 57.14%
New primary decisions (640) 45.00% 62.66% 40.16%
Both new decisive answers correct 4.06% 35.31% 3.75%

The V13 repeat differs on one of 960 labels from its original GPU run. Preserve these separate comparisons rather than mixing rows. Several older V13 suites regress relative to V12. GLiNER has a different architecture, native adapter and 512-token context. All 20 suites, older comparison including OpenJev, creator metrics.

These are adapted tests, not official leaderboard submissions. The public 231-item JevBench subset is not the full/private JevBench Score. Training and checkpoint selection exclude JevBench and sealed final cases; upstream exposure is unknown. Because older test scores have been observed across rounds, they are regression evidence. Numeric grading, calibrated confidence, real trading and unrestricted business or software decisions remain unvalidated.

CPU inference

Use Python 3.10+ in a virtual environment. Tested CPU dependencies are Torch 2.14.0+cpu and Transformers 5.17.0. The standalone wrapper uses explicit BF16 and rejects inputs over 4,096 tokens.

python -m pip install 'torch==2.14.0' --index-url https://download.pytorch.org/whl/cpu
python -m pip install 'transformers==5.17.0' 'huggingface-hub==1.33.0' 'safetensors==0.8.0'
python - <<'PY'
from huggingface_hub import hf_hub_download
for name in ['el_jevo.py', 'examples/ci-decisions.json']:
    hf_hub_download('DanielAlonsoC/el-jevo', name, revision='V17', local_dir='el-jevo')
PY
cd el-jevo
python el_jevo.py --revision V17 --input examples/ci-decisions.json --threads 4

Specify the version explicitly. For stock Transformers, use AutoModelForSequenceClassification.from_pretrained('DanielAlonsoC/el-jevo', revision='V17', subfolder='0.8b', dtype=torch.bfloat16). No remote code is required. Format inputs as Premise: {premise}\nHypothesis: {hypothesis}; class order is contradiction, entailment, neutral. The wrapper constructs candidate hypotheses and normalizes their entailment probabilities; these are not guaranteed calibrated estimates. CPU lacks the fast CUDA recurrent kernels, so long inputs can be slow. V17 standalone compatibility checks.

For Python usage after downloading the wrapper:

from el_jevo import ElJevo
model = ElJevo(revision="V17", threads=4)
print(model.nli("The build passed every test.", "The build passed its tests."))
# model.decide(payload) accepts examples/ci-decisions.json's typed input format.

The V17 wrapper defaults to V17, and the examples above explicitly pin it. Older immutable tags preserve their release-time runtime and documentation snapshots.

V17 training and provenance

V17 continues the V13 model on 26,020 training rows and 6,523 development rows, using the same corpus as the preceding retention experiments. One epoch completed 3,038 optimizer updates; development selected update 1,000, not the last checkpoint. Rank-16 LoRA covers all text linear layers, with a frozen native NLI classifier and learning rate 5e-6. Native-NLI and normalized-candidate KL retention terms both have weight 8, anchored only to V13's previously correct answers on old training rows. Incorrect reference answers remain eligible for correction from verified labels.

No new teacher data was purchased for V17. The merged model keeps the same parameter count and its classifier bytes match V13 exactly. CPU compatibility checks passed binary, choice, numeric and NLI forms, with exact probability parity against the evaluated adapter. These are compatibility checks, not a CPU quality or latency benchmark. Recipe and completion evidence · V17 CPU checks · Artifact manifest.

Earlier V12/V13 training recipes remain available as historical provenance: V12, V13. The corpus includes program-verified labels, reviewed normal-GLM language and MultiNLI retention examples. Earlier GLM 5.3 generation used OpenRouter/Venice. No OpenAI-generated answers were used as V17 training labels. JevBench and sealed final inputs were excluded from training and checkpoint selection; upstream exposure is unknown. No final benchmark selected a checkpoint.

Credit and licenses

Thank you to AlexWortega, creator of OpenJev, for the starting checkpoint and decision-model work. The original pin is 058a6c24911b46d908fbe23541390f8af3df3e4d, subfolder qwen3.5-0.8b-nli-v5, based on Qwen3.5-0.8B.

EL JEVO additions are MIT-licensed; preserve OpenJev attribution and Qwen Apache-2.0 obligations. LICENSE, NOTICE, Qwen license. MultiNLI sources have mixed licenses. The package contains merged runtime files, the inference wrapper, a toy example and aggregate evidence; training corpora, teacher weights, test cases, raw predictions and credentials are not distributed. Training repository.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for DanielAlonsoC/el-jevo

Finetuned
Qwen/Qwen3.5-4B
Finetuned
(5)
this model

Dataset used to train DanielAlonsoC/el-jevo