Instructions to use Jwuthrich/selfjev-4b with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use Jwuthrich/selfjev-4b with PEFT:
Task type is invalid.
- Notebooks
- Google Colab
- Kaggle
SelfJev-4B
Structured decisions from text. A 4B backbone. Your own GPU.
Route a customer message. Check a claim. Review an AI response. SelfJev answers questions over the options you supply and returns probabilities through a Python SDK or HTTP API. Its default engine shares the document computation across questions and scores candidates directly, without generating an answer token by token.
Quickstart · Results · Public datasets · Model size · Architecture · Latency · Research · Source code
| Ask for | Get back | Use it for |
|---|---|---|
| Yes / no | Probability of yes | Verification, guardrails, detection |
| One choice | Selected option and probabilities | Routing, intent, categorization |
| An ordered score | Score and probabilities over your levels | Review and grading |
| Every matching option | Selected set and per-option probabilities | Tags, checklists, multiple issues |
This repository contains the 230 MB LoRA adapter, not the entire model. The Qwen3.5-4B base weights download separately. The 4B model size refers to the backbone; the adapter stores the learned update.
Results
| Evaluation | Questions | SelfJev-4B | Jev API | SelfJev evidence |
|---|---|---|---|---|
| Text decisions | 1,991 | 95.7% | 97.2% | report |
| AI response review | 946 | 93.1% | 92.5% | report |
| Public-source tasks | 3,300 | 83.4% | 82.1% | report |
What these measure. Text decisions cover authored routing, policy, evidence and difficult language cases. AI response review covers verification, judging, scoring, guardrails and jailbreak detection. Public-source tasks cover 11 adapted datasets. Each score counts a question as correct only when the expected answer matches; for multiple selections, the entire set must match.
These are recorded results from the current TreeServer engine and matched Jev predictions, not scores from a newly merged vLLM release. The authored sets use AI-checked labels, not human ground truth. Both have informed research direction. The public-source set is a repeatedly reused development benchmark. A fresh independent evaluation remains necessary.
How does it compare with another open model?
On the same 1,991 text-decision questions, the recorded Eikos-4B run scored 92.8% using our task adapter and its letter-based readout. This is a comparison under this project's protocol, not a general ranking of open models. Eikos report
There is no SelfJev S1Bench result yet. Lev's published S1Bench scores cannot be compared with our authored-test scores. Comparing SelfJev, Lev, Reflex and other decision models requires the same dataset revision, item IDs, candidate sets, scoring rules and coverage.
Public datasets
83.4% across 3,300 questions. There are 300 questions per source, so the macro average over sources and the question-weighted average are equal. The 171 authored development questions are excluded here; the full 3,471-question development benchmark remains available in the reports.
| Source dataset | Task in our protocol | Training exposure | Questions | SelfJev | Jev |
|---|---|---|---|---|---|
| CLINC150 | Intent routing | held-out source | 300 | 95.0% | 94.3% |
| DBpedia | Topic classification | held-out source | 300 | 98.0% | 98.0% |
| TREC | Question type | held-out source | 300 | 91.3% | 94.0% |
| Emotion | Emotion classification | held-out source | 300 | 57.3% | 57.7% |
| BoolQ | Reading comprehension | held-out source | 300 | 87.7% | 90.7% |
| SST-2 | Sentiment | held-out source | 300 | 91.3% | 96.7% |
| Banking77 | Banking intent | seen source | 300 | 96.3% | 95.7% |
| AG News | News topic | seen source | 300 | 90.3% | 87.3% |
| TweetEval | Tweet sentiment | seen source | 300 | 67.3% | 64.3% |
| MNLI | Entailment (binary) | seen source | 300 | 95.0% | 92.7% |
| GoEmotions | Emotion tags (exact set) | seen source | 300 | 47.7% | 31.7% |
“Held-out source” means that dataset was excluded from SelfJev's task-specific training. “Seen source” means other rows from that source were used in training. Neither establishes absence from Qwen's pretraining.
These are adapted tasks, not official dataset leaderboard scores. For example, Banking77 uses eight candidate options per question, CLINC uses seven or eight, and DBpedia uses six. MNLI is recast as binary entailment. GoEmotions scores exact matching over six candidate tags. Our preprocessing, sampling and answer sets are part of the evaluation and must accompany the numbers.
Known development-data issues include shared MNLI premises across splits and near-duplicate authored policy texts. See data methodology. All 11 slices are shown, including weaker emotion and sentiment results.
Download chart data · SelfJev predictions · Matched Jev predictions
Accuracy and model size
| Model / recipe | Backbone size | Training | Accuracy | Evidence |
|---|---|---|---|---|
| Qwen3 reranker 0.6B | 0.6B | stock | 62.2% | report |
| Qwen3 reranker 4B | 4B | stock | 63.8% | report |
| Qwen3 reranker 8B | 8B | stock | 67.0% | report |
| Qwen3 + LoRA 0.6B | 0.6B | early fine-tune | 74.5% | report |
| Qwen3 + LoRA 4B | 4B | early fine-tune | 80.8% | report |
| Qwen3 + LoRA 8B | 8B | early fine-tune | 81.0% | report |
| SelfJev-4B | 4B | current release | 83.4% | report |
All points use the same 3,300 question IDs, labels and candidate sets. The stock models are rerankers mapped to our decision tasks. The earlier LoRA runs use older training recipes; SelfJev uses a different backbone generation and more training data. This shows the measured systems' trade-off between size and accuracy. It does not isolate the effect of parameter count or establish that a 4B model outperforms larger models generally.
Sizes are nominal backbone parameter counts, not checkpoint megabytes or runtime memory. Jev appears as a horizontal reference because its parameter count is undisclosed. We do not assign a guessed size to proprietary models.
AI response review
Breakdown of the 946-question authored suite, using the same current-engine predictions as the overview:
| Task | Questions | SelfJev | Jev |
|---|---|---|---|
| Factual verification | 165 | 87.3% | 86.7% |
| Response judging | 222 | 94.1% | 94.6% |
| Quality scoring | 200 | 92.5% | 93.0% |
| Policy guardrails | 170 | 96.5% | 92.9% |
| Jailbreak detection | 189 | 94.7% | 94.2% |
Labels were authored and checked with AI judges. Questions about one document are correlated; small differences should not be read as proof of superiority. This suite measures our defined answer choices and criteria, not unrestricted judgment quality. Dataset design · Predictions
Quickstart
On Linux with an NVIDIA GPU, Git and uv installed:
GIT_LFS_SKIP_SMUDGE=1 git clone https://github.com/Jwuthri/SelfJev.git
cd SelfJev
uv sync --frozen --no-dev --extra serve --extra gpu
uv run --no-sync python - <<'PYTHON'
from huggingface_hub import snapshot_download
snapshot_download(
"Jwuthrich/selfjev-4b",
local_dir="weights/selfjev_4b_hf",
allow_patterns=["adapter_model.safetensors", "adapter_config.json", "model.json"],
)
PYTHON
# Choose a secret for your own server.
export SELFJEV_API_KEYS="replace-with-your-long-random-secret"
uv run --no-sync selfjev serve \
--adapter weights/selfjev_4b_hf --host 127.0.0.1 --port 8000
In another terminal, use the same secret:
from selfjev import SelfJev, Noul, Choice
client = SelfJev(
base_url="http://127.0.0.1:8000",
api_key="replace-with-your-long-random-secret",
)
result = client.system_one(
state="I was charged twice. Please refund the duplicate payment.",
questions={
"refund": Noul("Does the customer want a refund?"),
"team": Choice("Which team should handle this?", {
"billing": "payments and refunds",
"support": "technical issues",
}),
},
)
print(result.nouls["refund"].noul)
print(result.choices["team"].choice)
The example shows the interface; it does not present fabricated model output. A generic Transformers text-generation or classification pipeline does not reproduce SelfJev's prompts, scoring rules or API. The server key is a secret you choose for your deployment, unrelated to Hugging Face credentials.
HTTP API · Self-hosting and AWS · Fine-tuning
Architecture
The default engine builds a shared-prefix tree: compute the document once, branch for each question, then branch for candidate answers. It reads yes/no evidence from the logits and converts those scores into the requested answer type. The server includes every candidate in the question, matching training.
The tree saves repeated computation at two levels: every question reuses the document, and every candidate reuses its question. In full-attention layers, the tree mask lets a branch read its ancestors but not sibling branches. In DeltaNet layers, branches inherit the recurrent state of their parent. Each candidate therefore sees its own document → question → candidate path, up to numerical differences between kernels. The diagram describes TreeServer; vLLM uses its own prefix-cache execution.
| Component | Current release |
|---|---|
| Backbone | Qwen3.5-4B |
| Base revision | 851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a |
| Fine-tuning | Rank-64 LoRA on attention and DeltaNet projections |
| Default serving | TreeServer, shared document and question computation |
| Alternative serving | vLLM with a merged checkpoint and prefix caching |
| Training examples | 79,943 non-test questions; texts up to 16K tokens |
| Training target | 50% checked labels + 50% stored Jev probabilities |
| Training schedule | One epoch, learning rate 2e-4 |
Jev probabilities served as a teacher signal; they did not decide the authored training labels. Authored labels were checked by an AI judge. Full lineage is recorded in model.json, with architecture detail in the research documentation.
The research path
The improvements came from several changes: verified difficult examples, the instruct backbone, explicit option lists, broader training coverage, tree-aware training and soft teacher targets. The chart is a history of complete systems, not an ablation assigning each gain to one change.
The research also includes smaller stock rerankers, 8B models, Jina, T5Gemma, custom architectures, distillation and RLCD. The full experiment ledger records results and dead ends; earlier implementations remain at the archive tag documented there. The checkpoint's original engine recorded a slightly different text-decision score; this card consistently uses the current engine for SelfJev's headline results.
Hardware and speed
A 24 GB NVIDIA GPU is a practical starting recommendation, not a measured minimum. Plan for 16–32 GB host RAM and at least 50 GB free disk. The adapter download size does not describe inference memory: the full backbone and runtime state must fit too.
The default engine avoids token-by-token answer generation and shares work across questions. Latency still depends on input length, question count, candidate count, GPU and concurrency. We have two sets of measurements below, with different models and engines. Current SelfJev-4B has no controlled NVIDIA latency sweep yet. Its CPU minimum is also unmeasured.
Earlier prototype on NVIDIA GPUs
| Text tokens | Questions | A10G | L40S | H100 |
|---|---|---|---|---|
| 8 | 1 | 55 ms | 36 ms | 22 ms |
| 512 | 1 | 136 ms | 55 ms | 30 ms |
| 2,048 | 1 | 361 ms | 121 ms | 47 ms |
| 4,096 | 1 | 698 ms | 228 ms | 82 ms |
| 8 | 16 | 401 ms | 135 ms | 58 ms |
| 512 | 16 | 496 ms | 163 ms | 76 ms |
| 2,048 | 16 | 808 ms | 268 ms | 120 ms |
| 4,096 | 16 | 1,263 ms | 424 ms | 189 ms |
These are server-side medians, excluding network time: ten warmed, successful requests per point, one request at a time, three answer options per question. All three sweeps use the archived tree_4b_combo Qwen3-4B model on vLLM in bf16. They were recorded in separate runs, not a simultaneous controlled hardware experiment. The plot uses a labeled logarithmic time axis; the table gives the actual milliseconds.
These measurements show what the earlier system achieved; they are not a latency promise for the current Qwen3.5 model or the merged release. A10G samples · L40S samples · H100 samples
Current SelfJev-4B on Apple Silicon
| Text tokens | Questions | M5 Pro / 48 GB |
|---|---|---|
| 8 | 1 | 688 ms |
| 8 | 16 | 6,786 ms |
| 512 | 1 | 2,018 ms |
| 512 | 16 | 8,599 ms |
This uses the current adapter merged into Qwen3.5-4B, the TreeServer engine, and PyTorch MPS in bf16 on a 20-core Apple M5 Pro GPU with 48 GB unified memory. Each cell is the median of ten timed calls after two warmups, with MPS synchronized before and after each call. Inputs are synthetic repeated text, with three answer options per question. Timing includes tokenization and model work; it excludes HTTP/network.
The MPS run uses the slower PyTorch recurrent fallback. It is a lower-level engine experiment, not a validated Mac server. Model and engine differences prevent a hardware-only comparison against the NVIDIA chart. The merged download also has not been benchmarked through vLLM using these workloads.
Download latency samples, configuration and medians · Latency CSV · Full speed research
Limits and reproducibility
- Probability is not a guarantee. The reported current-engine runs have no fitted calibration attached. Validate thresholds on application data before acting on confidence.
- Some tasks remain weak. Emotion labeling, fine-grained sentiment and selecting an exact set of labels are visibly harder than topic and intent classification.
- Benchmarks have a scope. AI-authored tests, reused development sets, sampled options and known overlap issues limit the conclusions. We make no S1Bench or general state-of-the-art claim.
- Engine versions matter. These accuracy results refer to the recorded current TreeServer runs. A merged checkpoint, quantization or another serving backend needs its own verification.
Every plotted value is regenerated from saved predictions or raw timing samples. The generator checks matched IDs, targets, question types and candidate sets for report-to-report comparisons, and recomputes latency medians from ten timed samples per cell. Chart data and source SHA-256 hashes · Generator · Architecture and latency plotting module · Card template
Place the downloadable generator, plotting module and template in scripts/docs/ of a SelfJev source checkout, then run:
uv run --no-project --with matplotlib==3.11.1 python scripts/docs/build_model_card.py
This builds figures from reports and the canonical data/all.jsonl.gz; it performs no model inference. Rebuild the canonical file using the data instructions if absent.
| File | Purpose |
|---|---|
adapter_model.safetensors |
Trained LoRA weights |
adapter_config.json |
PEFT configuration |
model.json |
Base revision, training provenance and original recorded scores |
assets/*.png, assets/*.svg |
Model-card figures, raster and vector |
assets/chart-data.json, assets/public-datasets.csv |
Underlying results and source provenance |
assets/latency-data.json, assets/latency-data.csv |
Selected raw timing samples, methodology and recalculated medians |
Adapter SHA-256: dfbf2834d883987893ec305a6093a345fd79c77f903f60f2f914cbe3b6058d1b.
The adapter uses the Qwen base model linked above. No separate adapter license is declared in this repository. Independent project; not affiliated with TypeSafe or Qwen.
- Downloads last month
- 42







