CareJev-Omni

A multimodal clinical decision classifier: give it a patient state (text, chest X-ray, rendered physiological signal, or heart-sound recording), a question, and the allowed answers; it returns a calibrated probability for each answer — no generated text. It implements the typed-decision interface (noul yes/no · choice · score) and is a drop-in variant of Jev-Omni: same prompt, same 256-slot decision head layout, same loader contract. Built on google/gemma-4-12B-it; trained only on open datasets prepared with PyHealth — no credentialed PHI.

What a System 1 model is, and why answer this way

Psychologists separate fast, intuitive judgement (System 1) from slow, step-by-step reasoning (System 2). TypeSafe AI uses the name System One model for models built around the first kind: the possible answers are decided in advance, and the model makes a single quick judgement over them with a probability attached, rather than writing an answer out word by word. CareJev-Omni is a System 1 model for clinical records, images and signals, and answering this way has its own advantages:

  • The answer is already structured. The question comes with its allowed answers and the model returns a probability for each one; there is no prose to read, parse or validate, and every answer is one of the options asked about.
  • Its confidence tracks how often it is right. When it says 80 %, it is correct about 80 % of the time: across all held-out test questions, stated confidence and actual accuracy differ by under 2 points on average (expected calibration error 0.018). Higher confidence means more likely correct: on the half of questions it is most sure about, accuracy is 0.897, against 0.741 over all questions (see the reliability diagrams below). A confidence threshold can therefore decide which answers to use as they are and which to pass to a slower process — a reasoning model or a person.
  • Position cannot steer it. The order in which the answers are listed does not change the result.
  • It is fast and predictable. One forward pass per question, with no generation loop and no variable-length output, and the same interface for text, X-rays, ECGs and heart sounds.

The framing also sets the limit: a fast judgement suits screening many questions and flagging the uncertain ones in research pipelines. It does not replace deliberate clinical reasoning, and this model is not validated for any decision about a real patient (see the notice below).

Research use only — not a medical device

This model is released for research purposes only. It is not a medical device, has not been reviewed or cleared by any regulator, and has not been validated on any live clinical population. Do not use it for diagnosis, triage, treatment, trial eligibility, resource allocation or any other decision about a real person. It is distributed "as is", without warranty of any kind, express or implied, and without any representation of fitness for a particular purpose; the authors accept no liability for any use of it or of its outputs. Anyone evaluating it does so at their own risk and remains responsible for complying with the licences below and with all applicable law and research-ethics requirements. Its outputs are probability estimates from an incomplete picture of a patient, carry the limitations listed under Limitations, and must never substitute for the judgement of a qualified clinician.

Results — held-out test split (patient-disjoint)

Every row is the same test items under an identical protocol for all systems: same prompts, same option lists, same split. base is the untouched google/gemma-4-12B-it scored by its next-token distribution over the option numbers (it placed ≥ 99 % of its mass on valid options). Jev-Omni, when present, is the published akhilaaa3/Jev-Omni checkpoint run through its own loader on the same items. majority is the majority-class rate. The best system in each row is in bold; bootstrap 95 % CIs (500 resamples) are given for the pooled figures. 6 sources measured here are not tabulated — a model given only the patient's sex and age answers them as well as this one, so per-source numbers there would not measure what they appear to measure; they remain in the totals and in the run artifacts.

By task family

15,629 held-out test questions.

Task family majority CareJev base Jev-Omni
12-lead ECG (rendered) (15) 0.737 0.792 0.587 0.529
Heart sounds (3) 0.694 0.700 0.429 0.404
Chest X-ray 0.490 0.870 0.694 0.571
Sleep staging (rendered PSG) 0.212 0.633 0.228 0.246
ICU outcomes (SUPPORT2) (4) 0.573 0.688 0.394 0.380
EHR demo (eICU / MIMIC-IV) (5) 0.465 0.460 0.279 0.242
Exam questions (MedMCQA) 0.303 0.728 0.715 0.708
Transcription specialty 0.264 0.445 0.417 0.421
all test items (pooled) 0.741 0.529 0.488
mean over sources 0.739 0.561 0.528

CareJev 95 % CI: pooled 0.741 (0.734–0.747); mean over sources 0.739 (0.728–0.750).

Ranking quality (AUROC, yes/no questions, positive = "Yes")

Task family CareJev base Jev-Omni
12-lead ECG (rendered) (15) 0.812 0.672 0.656
Heart sounds (3) 0.626 0.495 0.543
ICU outcomes (SUPPORT2) (4) 0.829 0.650 0.651
EHR demo (eICU / MIMIC-IV) (5) 0.650 0.571 0.508

Calibration (expected calibration error, 10 equal-width bins; lower is better)

Task family CareJev base Jev-Omni
12-lead ECG (rendered) (15) 0.062 0.331 0.197
Heart sounds (3) 0.148 0.472 0.250
Chest X-ray 0.044 0.268 0.187
Sleep staging (rendered PSG) 0.066 0.491 0.145
ICU outcomes (SUPPORT2) (4) 0.037 0.571 0.453
EHR demo (eICU / MIMIC-IV) (5) 0.095 0.607 0.408
Exam questions (MedMCQA) 0.037 0.241 0.132
Transcription specialty 0.040 0.559 0.499
all test items (pooled) 0.018 0.394 0.250
mean over sources 0.092 0.367 0.228

Reliability diagrams (confidence vs accuracy) per modality

Family rows weight each question by its test items (so accuracy is the pooled family accuracy); AUROC is the mean over the family's yes/no questions; the number in brackets is how many questions a family has (the datasets are credited under Data). The checkpoint was selected on the dev split only; the test split was never used for selection or calibration.

One-shot external benchmarks (never trained on, run once)

Every set below was converted to the typed-decision format and scored once, after the internal test, for CareJev, the untouched base model and the published Jev-Omni (through its own loader) on identical items. Nothing was tuned on them. They answer two different questions: the knowledge sets (MedQA, MMLU, PubMedQA, DecisionBench) ask what the fine-tune kept or lost of the base model's general ability; the shift sets (VQA-RAD, MedMNIST, CinC-2016, PTB-XL) ask whether the clinical training transfers to other hospitals, devices and question wordings.

Accuracy

benchmark majority CareJev base Jev-Omni
MedQA-USMLE (4-option) 0.277 0.738 0.687 0.701
MMLU medical (8 subjects) 0.304 0.821 0.821 0.812
PubMedQA (yes/no/maybe) 0.552 0.746 0.732 0.738
DecisionBench medium 0.365 0.792 0.761 0.867
DecisionBench hard 0.355 0.590 0.567 0.689
VQA-RAD closed (radiology) 0.530 0.733 0.753 0.781
MedMNIST Pneumonia (CXR) 0.630 0.602 0.876 0.836
MedMNIST Derma (7-class) 0.672 0.456 0.314 0.152
MedMNIST Blood (8-class) 0.178 0.368 0.412 0.344
MedMNIST OrganA (11-class CT) 0.174 0.288 0.334 0.310
PhysioNet/CinC 2016 heart sounds 0.800 0.617 0.200 0.200
PTB-XL ECG (5 tasks) 0.755 0.785 0.318 0.299
all items (pooled) 0.678 0.577 0.566
mean over benchmarks 0.667 0.503 0.495

DecisionBench's own scoring (state-macro accuracy · confidence − accuracy gap): CareJev medium 0.821 (-14 pp), hard 0.555 (+2 pp) · base medium 0.771 (+19 pp), hard 0.551 (+37 pp) · Jev-Omni medium 0.882 (+1 pp), hard 0.643 (+15 pp)

Calibration (ECE, lower is better)

benchmark CareJev base Jev-Omni
MedQA-USMLE (4-option) 0.054 0.209 0.107
MMLU medical (8 subjects) 0.038 0.161 0.100
PubMedQA (yes/no/maybe) 0.029 0.244 0.171
DecisionBench medium 0.146 0.198 0.030
DecisionBench hard 0.118 0.380 0.165
VQA-RAD closed (radiology) 0.067 0.213 0.083
MedMNIST Pneumonia (CXR) 0.183 0.106 0.033
MedMNIST Derma (7-class) 0.048 0.580 0.549
MedMNIST Blood (8-class) 0.110 0.529 0.372
MedMNIST OrganA (11-class CT) 0.068 0.481 0.277
PhysioNet/CinC 2016 heart sounds 0.076 0.608 0.300
PTB-XL ECG (5 tasks) 0.082 0.635 0.496
all items (pooled) 0.014 0.354 0.228
mean over benchmarks 0.084 0.430 0.292

PTB-XL: compare AUROC, not accuracy. We render the PTB-XL ECGs in the 12-lead paper layout this model was trained on. That layout drags the base model's accuracy down (hypertrophy 0.176, against 0.856 on a plainer strip layout) without changing how well it ranks cases, so accuracy would overstate the margin. AUROC does not depend on a decision threshold, and CareJev has the highest AUROC on all 5 tasks, against either layout:

PTB-XL task (AUROC) CareJev base Jev-Omni base, strip layout
conduction disturbance 0.665 0.603 0.634 0.605
hypertrophy 0.664 0.442 0.447 0.603
myocardial infarction 0.781 0.663 0.653 0.670
normal ECG 0.844 0.594 0.598 0.678
ST/T changes 0.820 0.399 0.474 0.535

Caveats that belong next to those numbers: MedQA and MMLU are almost certainly in the base model's pre-training; MedMNIST-Pneumonia draws on the same paediatric CXR collection as the "viral pneumonia" class of our COVID-19 Radiography training data, so its number may be optimistic; and on CinC-2016 heart sounds no model beats the trivial baseline: answering "normal" to everything scores 0.800, this model scores 0.617 with an AUROC of 0.542, and the base model and Jev-Omni score 0.200 because they answer "abnormal" to almost every recording. Being better than the other two there is not the same as being useful: the honest reading is that the heart-sound pathway does not transfer to this collection, and it is the open item for the next run.

Further analyses

Post-hoc calibration

Fitted on the dev split: a global temperature (T = 1.036), with per-source temperature / yes-bias terms where they helped. Not shipped. The fit was validated on the single-prompt path (test NLL 0.6377 to 0.6284), but the released path averages over option rotations, which already smooths the probabilities: it measures an expected calibration error of 0.018 without any post-hoc term. Shipping a correction that was never validated on the path we release would make the tables describe a configuration nobody runs, so the loader returns the raw head probabilities and every number on this card is measured that way. The fit is kept with the run artifacts for the record.

before after (fitted, not shipped)
dev NLL (fit set) 0.6740 0.6738
test NLL 0.6377 0.6284
test ECE 0.0218 0.0096
test mean-source NLL 0.6127 0.6084
test accuracy 0.7375 0.7389

Selective prediction (does confidence rank errors?)

Accuracy when only the most confident share of test questions is answered; a model whose confidence is informative gains accuracy as coverage falls.

15,629 test questions (text 5,847, image 8,835, audio 947).

group acc @ 50 % acc @ 80 % acc @ 100 %
all 0.897 0.815 0.741
text 0.868 0.752 0.670
image 0.921 0.846 0.791
audio 0.829 0.748 0.711

Beyond accuracy (balanced accuracy and score questions)

On skewed questions accuracy can sit at the majority rate while the model still separates the classes; balanced accuracy (the mean of per-class recall) shows that. For score questions the interface returns the expected level E[level] = Σ i·p_i.

Task family CareJev base Jev-Omni
12-lead ECG (rendered) (15) 0.630 0.570 0.546
Heart sounds (3) 0.515 0.501 0.505
Chest X-ray 0.798 0.457 0.434
Sleep staging (rendered PSG) 0.651 0.225 0.255
ICU outcomes (SUPPORT2) (4) 0.590 0.434 0.423
EHR demo (eICU / MIMIC-IV) (5) 0.222 0.146 0.130
Exam questions (MedMCQA) 0.725 0.710 0.695
Transcription specialty 0.373 0.555 0.514

Score questions, mean absolute error of the expected level: eicu_demo_los 1.40 levels (majority 1.77); mimic4_demo_los 2.31 levels (majority 2.38); support2_los 1.53 levels (majority 1.57).

Sub-groups (SUPPORT2, as recorded in the state)

2,739 SUPPORT2 test questions (in-hospital death, 2- and 6-month survival); groups with fewer than 30 questions are not shown.

group accuracy AUROC ECE
sex: female 0.758 0.841 0.019
sex: male 0.790 0.869 0.030
age: <50 0.803 0.893 0.043
age: 50-64 0.791 0.862 0.034
age: 65-79 0.747 0.828 0.014
age: 80+ 0.780 0.863 0.057
race: black 0.803 0.873 0.036
race: hispanic 0.758 0.892 0.138
race: other 0.833 0.863 0.108
race: white 0.769 0.850 0.025

Latency of the released path

path questions ms / question input tokens / s
noul (Yes/No), one prompt 40 307 925
score, 10 levels, released path (8 rotations, one batch) 40 2287 154
score, 10 levels, single prompt (permutations=0) 40 338 1042
prefix cache (opt-in), 3-4 questions per state, single prompt each 40 183 -

Measured on NVIDIA GB10.

Data

CareJev-Omni was trained only on publicly available, open-access medical datasets: rendered 12-lead ECGs, chest X-rays, heart-sound recordings, sleep recordings, ICU outcome records, de-identified EHR demo data, medical exam questions and clinical transcriptions. No credentialed data and no protected health information were used, and evaluation uses a patient-disjoint held-out split.

Attribution: PhysioNet/CinC Challenge 2020, CirCor DigiScope, Sleep-EDF, eICU-CRD demo and MIMIC-IV demo (PhysioNet); COVID-19 Radiography Database (CC BY 4.0); BMD-HS (CC BY 4.0); SUPPORT2 (Vanderbilt / UCI); MedMCQA (MIT); Medical Transcriptions (CC0). Each dataset's own licence and terms continue to apply.

Quick start

pip install -r requirements.txt          # torch, torchvision, transformers==5.17.0, peft, ...
# audio input also needs ffmpeg on PATH

torchvision is required even for text-only use: transformers' Gemma 4 image processor imports it at module load. This repo is self-contained: the frozen base weights, the adapter applied on top at load and the decision head. No network access is needed after download. The low-rank delta is deliberately not folded into the weights: merging it into bf16 leaves accuracy and NLL intact but costs about half of this model's calibration and flips 3.4 % of answers (measured, 300 test items), and calibration is the property this card is about. python example.py (shipped in this repo) prints one answer of each question type.

Every file's SHA-256 is in sha256.json; to check a download:

python -c "import json,hashlib,pathlib;[print(k, hashlib.sha256(pathlib.Path(k).read_bytes()).hexdigest()==v) for k,v in json.load(open('sha256.json')).items()]"
from huggingface_hub import snapshot_download
import sys; sys.path.insert(0, snapshot_download("nevermindai/carejev-omni", allow_patterns=["*.py"]))
from carejev_omni_loader import load_carejev_omni   # same contract as Jev-Omni's load_jev_omni

clf = load_carejev_omni("nevermindai/carejev-omni")
print(clf.predict(
    state="71-year-old male. Admission type: Emergency. ... DAY-3 PHYSIOLOGY: mean arterial pressure 55 mmHg; ...",
    question="Will this patient die before being discharged from this hospitalization?",
    options=["Yes", "No"],
))
# {'prediction': 'Yes', 'prediction_index': 0, 'confidence': 0.71, 'probabilities': {'Yes': 0.71, 'No': 0.29}}

The inference code ships in this repo (carejev_omni_loader.py and the carejev_omni/ package it imports), so pip install -r requirements.txt is all a fresh environment needs.

How to use

Once loaded, clf.predict(...) answers one question about one patient state, and clf.predict_many(...) answers several questions about the same state.

Inputs

argument type meaning
state str what is known about the patient, written as text (notes, values, context)
question str the question to answer
options list[str] the allowed answers: 2 to 256, all different
media path, optional an image (chest X-ray, rendered 12-lead ECG, spectrogram) or a heart-sound recording (.wav)
modality "text" · "image" · "audio" what media is; default "text" (no media)
qtype "noul" · "choice" · "score", optional inferred when omitted (["Yes", "No"] is a yes/no question, anything else is a choice); pass "score" when the options are ordered levels
permutations None or 0 None (default) averages over rotations of the options, so their order cannot change the answer; 0 scores one prompt in your order (faster)

Output: a plain, JSON-serialisable dict:

{
  "prediction": "Yes",
  "prediction_index": 0,
  "confidence": 0.71,
  "probabilities": {"Yes": 0.71, "No": 0.29}
}

probabilities holds every option, in the order you passed them, summing to 1. prediction is the most probable option, prediction_index its position in your list and confidence its probability. A threshold on confidence is how you decide which answers to use as they are and which to pass on for review.

The same structure comes back for every question type (values below are illustrative). Choice:

{
  "prediction": "Sepsis",
  "prediction_index": 0,
  "confidence": 0.82,
  "probabilities": {"Sepsis": 0.82, "Heart failure": 0.09, "Pulmonary embolism": 0.06,
                    "Diabetic ketoacidosis": 0.03}
}

Score (ordered levels): the most probable level, plus every level's probability:

{
  "prediction": "8 to 14 days",
  "prediction_index": 8,
  "confidence": 0.34,
  "probabilities": {"Less than 1 day": 0.0, "1 day": 0.01, "2 days": 0.02, "3 days": 0.04,
                    "4 days": 0.06, "5 days": 0.08, "6 days": 0.09, "7 days": 0.11,
                    "8 to 14 days": 0.34, "More than 14 days": 0.25}
}

predict_many: a list of the same dicts, one per question, in question order:

[
  {"prediction": "No", "prediction_index": 1, "confidence": 0.77, "probabilities": {"Yes": 0.23, "No": 0.77}},
  {"prediction": "Yes", "prediction_index": 0, "confidence": 0.58, "probabilities": {"Yes": 0.58, "No": 0.42}}
]

Examples

state = "PATIENT: 68-year-old male. ... DAY-3 PHYSIOLOGY: mean arterial pressure 72 mmHg; ..."

# yes / no: read the probability of "Yes"
out = clf.predict(state=state, options=["Yes", "No"],
                  question="Will this patient die before being discharged from this hospitalization?")
p_yes = out["probabilities"]["Yes"]

# choice: one of several options
out = clf.predict(state=state, question="Which single problem best explains this presentation?",
                  options=["Sepsis", "Heart failure", "Pulmonary embolism", "Diabetic ketoacidosis"])
answer, confidence = out["prediction"], out["confidence"]

# score: ordered levels; the expected level is the probability-weighted position
levels = ["Less than 1 day", "1 day", "2 days", "3 days", "4 days", "5 days", "6 days", "7 days",
          "8 to 14 days", "More than 14 days"]
out = clf.predict(state=state, question="How long will this patient's hospital stay be in total?",
                  options=levels, qtype="score")
expected_level = sum(i * out["probabilities"][o] for i, o in enumerate(levels))

# an image (here a chest X-ray) or a heart-sound recording (audio needs ffmpeg on PATH)
clf.predict(state="Frontal chest radiograph.", question="What does this chest X-ray most likely show?",
            options=["COVID-19", "Normal", "Lung opacity", "Viral pneumonia"],
            media="cxr.png", modality="image")
clf.predict(state="Phonocardiogram: digital stethoscope recording of heart sounds.",
            question="Is a heart murmur audible in this recording?", options=["Yes", "No"],
            media="pcg.wav", modality="audio")

# several questions about the same state: a list of the same dicts, in question order
outs = clf.predict_many(state=state, questions=[
    ("Will this patient die before being discharged from this hospitalization?", ["Yes", "No"]),
    ("Will this patient be alive 6 months (180 days) after study entry?", ["Yes", "No"]),
])

Questions phrased like the training data work best: the examples above use that wording, and the model card's task families show what it was trained on. predict_many(..., use_prefix_cache=True) encodes the state once for all the questions; it is faster but scores each question with a single prompt in your option order, like permutations=0. Every number on this card uses the default path.

Limitations

  • Trained on open, mostly small datasets; SUPPORT2 (1989–94, five US hospitals) dominates the text portion. Not validated on any live clinical population.
  • Signal "images" are our own renderings; performance depends on that rendering.
  • Heart-sound questions whose label age and sex alone predict as well as this model (bmdhs_ar, bmdhs_as, bmdhs_mr, bmdhs_normal) are not reported: a per-source number there would not measure the recording. The heart-sound rows that are tabulated (bmdhs_ms, circor_murmur, circor_outcome) do not have that shortcut, and on the external CinC-2016 collection no model beats the trivial answer.
  • Video is supported by the architecture but received no training data.
  • eICU / MIMIC-IV demo sources are tiny (≤ 100 patients); their numbers are noisy.
  • Calibration is reported on the internal test split only; distribution shift will degrade it.
  • For a single structured-tabular task, a tabular model is at least as good. On the SUPPORT2 outcomes, a gradient-boosted model fitted on the raw numeric columns available at study entry is level with this model on the three binary outcomes (AUROC within 0.01, accuracy within 0.01) and clearly better on the length-of-stay score question, with slightly lower NLL on all four. Per-source classical models also win on a few other sources (notably the tiny eICU demo tasks and CirCor murmur). This model is for the case where the evidence is mixed — prose, rendered signals, sounds — and one interface has to cover all of it, not for replacing a well-fitted tabular risk score.

Relation to other work

Independent open model implementing the typed-decision interface described by TypeSafe AI's Jev; not affiliated with or trained on outputs of TypeSafe AI, Jev, Jev-Omni, or Google. Architecture and prompt follow Jev-Omni for drop-in compatibility.

Evaluation protocol

  • Splits are patient-disjoint and were fixed before any model saw the data.
  • The test split is touched once per model, after training; selection uses dev only.
  • Baselines share the exact prompt text and option lists (the base model's answer is read from its next-token distribution over the option numbers).
  • Published comparisons are to the untouched base model and to the other released model implementing this interface. Non-neural baselines (TF-IDF / tabular / pixel / MFCC models fitted per source) were measured and are kept with the run artifacts, but are not part of the published tables: they are bespoke per task and per modality, so they do not answer the question this card is about, which is what one general model does across every source. See Limitations for where such a model is the better tool.
  • Intervals: non-parametric bootstrap over test examples, 500 resamples, percentile 95 % CI.
  • Calibration terms are fitted on dev and reported on test before/after.
  • Every number on this card is rendered directly from the evaluation outputs of this checkpoint.

Licence

The base weights are Gemma 4, published by Google under Apache-2.0 (licence). Our own contributions — decision head, loader code and this card — are Apache-2.0 as well, so the whole release is Apache-2.0. Dataset licences are listed above and remain separate. Both licences disclaim warranties and limit liability; nothing here grants any additional warranty or indemnity, and the research-use notice at the top of this card applies to every part of the release. Card generated 2026-10-04 from run artifacts.

Downloads last month
72
Safetensors
Model size
12B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for nevermindai/carejev-omni

Finetuned
(194)
this model

Space using nevermindai/carejev-omni 1