Instructions to use DanielAlonsoC/el-jevo with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use DanielAlonsoC/el-jevo with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-classification", model="DanielAlonsoC/el-jevo")# pip install -U transformers accelerate # Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("DanielAlonsoC/el-jevo", device_map="auto") - Notebooks
- Google Colab
- Kaggle
EL JEVO 0.8B — V17
EL JEVO is a small AI model for judging answers and making decisions from text. You give it a question, the relevant evidence or rules, and a set of possible answers. It evaluates how well each answer is supported and returns a structured decision with a score for each candidate. It runs locally, including on CPU.
EL JEVO 0.8B is a further fine-tune of OpenJev by AlexWortega, built on Qwen3.5-0.8B. The project experiments with additional training for applying written rules, interpreting facts and comparing answers while retaining the original model's ability to judge whether a claim follows from a passage.
| Task | What you supply | What it returns |
|---|---|---|
Yes/no (noul) |
Evidence and a question with yes/no answers | A selected answer and scores for yes and no |
Multiple choice (choice) |
Evidence, a question and described answer options | A selected option and scores for the candidates |
Numeric scale (score) |
Evidence, a written rubric and allowed numeric values | A selected value, candidate scores and an expected score |
| Natural language inference (NLI) | A passage and a claim about it | Entailment (the claim follows), contradiction (it conflicts), or neutral (the passage leaves it undetermined) |
For example, an application could ask EL JEVO to check a proposed action against a policy. Here is an illustrative task, with an expected answer that follows from the stated rule:
Evidence: A build may be released only when every test passes.
This build has 12 passing tests and 1 failing test.
Question: Can this build be released under the policy?
Possible answers: yes / no
Expected answer: no
The application supplies the evidence and answer choices. EL JEVO can serve as a local evaluator for candidate answers or proposed workflow actions; its scores express the model's assessment and are not guaranteed calibrated probabilities of correctness.
V17 is the current recommended EL JEVO release and the default checkpoint on main. Use revision='V17', subfolder 0.8b, to pin this release.
The complete merged BF16 checkpoint has 852,988,992 parameters (nominal 0.8B, including inherited unused vision weights), approximately 1.7 GB of stored weights, and runs on CPU. Text inference is evaluated; a GPU is optional. Detailed V17 checkpoint card.
Versions
| Revision | Role |
|---|---|
| V17 | Current recommended release; default weights on main |
| V13 | Historical predecessor and V17 training parent |
| V12 | Historical broader-retention alternative; some older scores differ |
| V10 | Historical older-rule retention alternative |
| V9 | Original published reference |
These are version tags of EL JEVO 0.8B, each stored under 0.8b/. V17 continues V13. The older cards and comparisons document earlier experiments; V13 is no longer the current recommended release. V17's recommendation reflects its development-selected checkpoint, modest gains on the sampled structural curriculum, and measured retention. It does not mean V17 wins every suite or generally beats every other model.
Evaluation
V17 compared with its V13 predecessor
The fixed V13, V16 and V17 checkpoints ran on 35,580 identical cases across 22 suites, on one RTX PRO 4500 Blackwell with the same BF16, context and batching settings. Development selected V17 at update 1,000 before the final confirmation. The new confirmation contains 445 roots, 1,335 total views and 86 held-out structural compositions.
| V17 promotion confirmation | V13 | V16 experiment | V17 |
|---|---|---|---|
| Primary decisions (890) | 62.13% | 62.58% | 63.82% |
| Both decisive answers correct (445 roots) | 33.93% | 34.83% | 37.53% |
| Invariant views (445) | 62.70% | 61.80% | 62.70% |
| Public JevBench (231) | 71.00% | 70.56% | 71.00% |
V17's primary gain over V13 is +1.69 percentage points: exploratory paired 95% interval [+0.22, +3.15] by root and [+0.32, +3.24] by structural composition. Both-correct decisive pairs improve by 3.60 points. Public JevBench remains 164/231 (71.00%). All 21 older suites stay within the preset two-percentage-point accuracy or 3% relative numeric-MAE regression limits; passing those limits allows small regressions. One training seed and shared language themes limit broader conclusions.
| Mechanism | Primary decisions | V13 | V17 |
|---|---|---|---|
| bounded uncertainty | 230 | 61.74% | 63.04% |
| causal interventions | 240 | 48.33% | 48.75% |
| joint allocation | 240 | 67.92% | 68.75% |
| quantified collections | 180 | 73.33% | 78.33% |
Causal intervention remains difficult, at 48.75% in this confirmation. Ordered scoring is measured in older suites; the four new mechanisms use binary and choice questions.
| Older regression suite | Cases | V13 | V17 |
|---|---|---|---|
| Public NLI sample | 396 | 64.39% | 64.39% |
| MNLI holdout | 120 | 88.33% | 88.33% |
| Operational probes | 324 | 87.04% | 87.04% |
| LLMBar | 838 | 52.98% | 53.34% |
| JudgeBench | 1240 | 59.03% | 59.03% |
| ContractNLI | 2091 | 61.36% | 61.07% |
| HelpSteer2 argmax numeric MAE ↓ | 2076 | 1.076 | 1.075 |
All 22 suites and paired intervals · Machine-readable promotion results · Development selection · Training experiment
V17 compared with original OpenJev 0.8B
A later completed inference-only follow-up, executed during the V18 round, compares the same published V17 weights with the original upstream OpenJev on the older V13 structural holdout. Both used the same GPU, BF16, context and batching. This is a separate 640-decision comparison; its scores are not mixed with the 890-decision confirmation above.
| Matched structural follow-up | Original OpenJev 0.8B | V17 |
|---|---|---|
| Primary decisions (640) | 281/640 — 43.91% | 405/640 — 63.28% |
| Both decisive answers correct (320 roots) | 13/320 — 4.06% | 118/320 — 36.88% |
| Invariant views (320) | 140/320 — 43.75% | 206/320 — 64.38% |
The primary gain is +19.38 percentage points, with a paired root-bootstrap 95% interval [+15.31, +23.28]. The four mechanism families and language themes are related to EL JEVO training. This older observed structural test is regression evidence; the result is a targeted task gain, not independent source transfer or universal superiority. Other external peers have not been evaluated in this follow-up. Complete comparison and boundaries · Checkpoint pins and aggregate evidence.
Historical comparisons with other small models
The full earlier comparison remains available: 20 checkpoints on 24,270 identical inputs, including OpenJev, Laya, Julia, Open-JEV DeBERTa, Lumma, Kev, Tiny-Jev, FineCat, EttinX, ModernCE, GLiNER2.5-Decide and EL JEVO V9–V13. V17 was not evaluated in that 20-checkpoint campaign. Its separate comparisons are reported above.
Show the earlier 20-checkpoint table and its evaluation limits
Historical 20-checkpoint comparison (V17 is not in this table)
20 fixed checkpoints, 24,270 identical cases across 16 suites. The upstream OpenJev 0.8B is the primary baseline. The table also includes Laya (three variants), Julia, Open-JEV DeBERTa, Lumma (two sizes), Kev (two sizes), Tiny-Jev, FineCat, EttinX, ModernCE (two sizes), GLiNER2.5-Decide and EL JEVO V9/V10/V12/V13. Every prediction file and input subset was verified before scoring.
| Model | JevBench (231) | NLI (396) | V10 rules (2,880) | LLMBar (838) | JudgeBench (1,240) | ContractNLI (2,091) | Prediction coverage |
|---|---|---|---|---|---|---|---|
| EL JEVO V13 (historical) | 71.00% | 64.39% | 41.56% | 52.98% | 59.03% | 61.36% | 100.00% |
| OpenJev 0.8B | 71.43% | 66.92% | 34.58% | 49.05% | 56.61% | 60.98% | 100.00% |
| EL JEVO V12 | 71.86% | 65.15% | 41.22% | 55.61% | 59.76% | 61.65% | 100.00% |
| EL JEVO V9 | 71.43% | 65.15% | 39.97% | 49.76% | 57.26% | 61.93% | 100.00% |
| EL JEVO V10 | 70.56% | 64.39% | 37.88% | 50.36% | 56.85% | 62.41% | 100.00% |
| laya | 48.05% | 63.13% | 33.26% | 40.10% | 0.08% | 1.24% | 81.11% |
| laya-typed-decisions | 50.22% | 60.61% | 32.43% | 49.76% | 20.81% | 5.93% | 88.54% |
| laya-multilingual | 45.45% | 55.81% | 33.61% | 50.72% | 51.61% | 44.57% | 99.50% |
| julia-1 | 44.59% | 32.58% | 33.33% | 47.61% | 49.11% | 44.43% | 99.50% |
| open-jev-deberta-v3-large | 48.48% | 33.08% | 33.78% | 27.80% | 0.00% | 0.00% | 76.63% |
| lumma-fev-0.6b | 46.75% | 35.86% | 32.64% | 49.28% | 49.60% | 37.35% | 99.07% |
| lumma-fev-0.1b | 37.66% | 33.84% | 34.65% | 49.28% | 34.84% | 13.15% | 92.17% |
| kev-0.8b | 63.20% | 44.19% | 35.73% | 54.42% | 54.76% | 48.59% | 99.09% |
| kev-0.5b | 48.92% | 40.66% | 33.16% | 52.03% | 49.60% | 38.02% | 99.16% |
| tiny-jev | 60.61% | 45.20% | 34.86% | 56.92% | 53.79% | 41.22% | 98.95% |
| finecat-nli-xxs | 47.19% | 57.07% | 33.33% | 53.46% | 50.40% | 20.76% | 99.30% |
| ettinx-nli-xxs | 47.62% | 43.43% | 33.12% | 58.11% | 52.42% | 23.72% | 99.30% |
| modernce-base-nli | 46.75% | 56.82% | 33.47% | 49.52% | 44.52% | 40.03% | 95.88% |
| modernce-large-nli | 50.65% | 65.40% | 33.78% | 42.72% | 46.45% | 48.78% | 95.88% |
| GLiNER2.5-Decide (512) | 57.14% | 34.34% | 34.06% | 40.10% | 49.35% | 44.62% | 100.00% |
On these saved runs, V13 scores 164/231 on public JevBench versus OpenJev's 165/231, and 64.39% on the NLI sample versus 66.92%. V13's V10-rule score is 41.56% versus 34.58%. This is a targeted rule-task gain; it does not establish general superiority over OpenJev or every peer.
Accuracy uses all cases; unsupported inputs count as wrong. The report also provides the same accepted subset for every model. Prediction coverage does not imply full evidence coverage: GLiNER uses a 512-token cap, and native peers have different input limits. These are saved quality measurements on identical inputs from separate inference runs (RTX 4090, RTX PRO 4500 Blackwell and A40), not one hardware-controlled run or a latency ranking. The V10 rules are older regression tests with generator families related to EL JEVO training.
All suites, common support and answer types · JSON · CSV · Original 17-model report
At the time of this historical table, the V13 structural report included GLiNER2.5-Decide but lacked original OpenJev. The later V17 versus OpenJev follow-up completes that upstream comparison using published V17 on the same structural inputs. The other external peers still lack completed structural-holdout results.
Show historical V12/V13 structural comparisons
Historical V12/V13 structural comparisons
| Same matched run: new structural holdout | V9 | V13 |
|---|---|---|
| Primary decisions (640) | 44.69% | 62.81% |
| Invariant views (320) | 44.06% | 61.88% |
| Both decisive answers correct (320 pairs) | 3.44% | 35.31% |
The paired primary gain is 18.125 percentage points, with a 95% root-cluster interval 14.375–22.031 points. Both fixed models used identical BF16/context/batching on the same RTX PRO 4500 Blackwell worker. The cases contain 320 roots and 83 held-out structural compositions in four mechanisms; language themes are shared across splits. This supports a targeted curriculum gain, not general decision-making superiority. Direct comparison.
| Original 20-suite comparison | V12 | V13 | GLiNER2.5-Decide |
|---|---|---|---|
| JevBench public (231) | 71.86% | 71.00% | 57.14% |
| New primary decisions (640) | 45.00% | 62.66% | 40.16% |
| Both new decisive answers correct | 4.06% | 35.31% | 3.75% |
The V13 repeat differs on one of 960 labels from its original GPU run. Preserve these separate comparisons rather than mixing rows. Several older V13 suites regress relative to V12. GLiNER has a different architecture, native adapter and 512-token context. All 20 suites, older comparison including OpenJev, creator metrics.
These are adapted tests, not official leaderboard submissions. The public 231-item JevBench subset is not the full/private JevBench Score. Training and checkpoint selection exclude JevBench and sealed final cases; upstream exposure is unknown. Because older test scores have been observed across rounds, they are regression evidence. Numeric grading, calibrated confidence, real trading and unrestricted business or software decisions remain unvalidated.
CPU inference
Use Python 3.10+ in a virtual environment. Tested CPU dependencies are Torch 2.14.0+cpu and Transformers 5.17.0. The standalone wrapper uses explicit BF16 and rejects inputs over 4,096 tokens.
python -m pip install 'torch==2.14.0' --index-url https://download.pytorch.org/whl/cpu
python -m pip install 'transformers==5.17.0' 'huggingface-hub==1.33.0' 'safetensors==0.8.0'
python - <<'PY'
from huggingface_hub import hf_hub_download
for name in ['el_jevo.py', 'examples/ci-decisions.json']:
hf_hub_download('DanielAlonsoC/el-jevo', name, revision='V17', local_dir='el-jevo')
PY
cd el-jevo
python el_jevo.py --revision V17 --input examples/ci-decisions.json --threads 4
Specify the version explicitly. For stock Transformers, use AutoModelForSequenceClassification.from_pretrained('DanielAlonsoC/el-jevo', revision='V17', subfolder='0.8b', dtype=torch.bfloat16). No remote code is required. Format inputs as Premise: {premise}\nHypothesis: {hypothesis}; class order is contradiction, entailment, neutral. The wrapper constructs candidate hypotheses and normalizes their entailment probabilities; these are not guaranteed calibrated estimates. CPU lacks the fast CUDA recurrent kernels, so long inputs can be slow. V17 standalone compatibility checks.
For Python usage after downloading the wrapper:
from el_jevo import ElJevo
model = ElJevo(revision="V17", threads=4)
print(model.nli("The build passed every test.", "The build passed its tests."))
# model.decide(payload) accepts examples/ci-decisions.json's typed input format.
The V17 wrapper defaults to V17, and the examples above explicitly pin it. Older immutable tags preserve their release-time runtime and documentation snapshots.
V17 training and provenance
V17 continues the V13 model on 26,020 training rows and 6,523 development rows, using the same corpus as the preceding retention experiments. One epoch completed 3,038 optimizer updates; development selected update 1,000, not the last checkpoint. Rank-16 LoRA covers all text linear layers, with a frozen native NLI classifier and learning rate 5e-6. Native-NLI and normalized-candidate KL retention terms both have weight 8, anchored only to V13's previously correct answers on old training rows. Incorrect reference answers remain eligible for correction from verified labels.
No new teacher data was purchased for V17. The merged model keeps the same parameter count and its classifier bytes match V13 exactly. CPU compatibility checks passed binary, choice, numeric and NLI forms, with exact probability parity against the evaluated adapter. These are compatibility checks, not a CPU quality or latency benchmark. Recipe and completion evidence · V17 CPU checks · Artifact manifest.
Earlier V12/V13 training recipes remain available as historical provenance: V12, V13. The corpus includes program-verified labels, reviewed normal-GLM language and MultiNLI retention examples. Earlier GLM 5.3 generation used OpenRouter/Venice. No OpenAI-generated answers were used as V17 training labels. JevBench and sealed final inputs were excluded from training and checkpoint selection; upstream exposure is unknown. No final benchmark selected a checkpoint.
Credit and licenses
Thank you to AlexWortega, creator of OpenJev, for the starting checkpoint and decision-model work. The original pin is 058a6c24911b46d908fbe23541390f8af3df3e4d, subfolder qwen3.5-0.8b-nli-v5, based on Qwen3.5-0.8B.
EL JEVO additions are MIT-licensed; preserve OpenJev attribution and Qwen Apache-2.0 obligations. LICENSE, NOTICE, Qwen license. MultiNLI sources have mixed licenses. The package contains merged runtime files, the inference wrapper, a toy example and aggregate evidence; training corpora, teacher weights, test cases, raw predictions and credentials are not distributed. Training repository.