Instructions to use nevermindai/carejev-omni with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use nevermindai/carejev-omni with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="nevermindai/carejev-omni") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("nevermindai/carejev-omni") model = AutoModelForMultimodalLM.from_pretrained("nevermindai/carejev-omni", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use nevermindai/carejev-omni with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "nevermindai/carejev-omni" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "nevermindai/carejev-omni", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/nevermindai/carejev-omni
- SGLang
How to use nevermindai/carejev-omni with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "nevermindai/carejev-omni" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "nevermindai/carejev-omni", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "nevermindai/carejev-omni" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "nevermindai/carejev-omni", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Docker Model Runner
How to use nevermindai/carejev-omni with Docker Model Runner:
docker model run hf.co/nevermindai/carejev-omni
CareJev-Omni
A multimodal clinical decision classifier: give it a patient state (text, chest X-ray, rendered physiological signal, or heart-sound recording), a question, and the allowed answers; it returns a calibrated probability for each answer — no generated text. It implements the typed-decision interface (noul yes/no · choice · score) and is a drop-in variant of Jev-Omni: same prompt, same 256-slot decision head layout, same loader contract. Built on google/gemma-4-12B-it; trained only on open datasets prepared with PyHealth — no credentialed PHI.
What a System 1 model is, and why answer this way
Psychologists separate fast, intuitive judgement (System 1) from slow, step-by-step reasoning (System 2). TypeSafe AI uses the name System One model for models built around the first kind: the possible answers are decided in advance, and the model makes a single quick judgement over them with a probability attached, rather than writing an answer out word by word. CareJev-Omni is a System 1 model for clinical records, images and signals, and answering this way has its own advantages:
- The answer is already structured. The question comes with its allowed answers and the model returns a probability for each one; there is no prose to read, parse or validate, and every answer is one of the options asked about.
- Its confidence tracks how often it is right. When it says 80 %, it is correct about 80 % of the time: across all held-out test questions, stated confidence and actual accuracy differ by under 2 points on average (expected calibration error 0.018). Higher confidence means more likely correct: on the half of questions it is most sure about, accuracy is 0.897, against 0.741 over all questions (see the reliability diagrams below). A confidence threshold can therefore decide which answers to use as they are and which to pass to a slower process — a reasoning model or a person.
- Position cannot steer it. The order in which the answers are listed does not change the result.
- It is fast and predictable. One forward pass per question, with no generation loop and no variable-length output, and the same interface for text, X-rays, ECGs and heart sounds.
The framing also sets the limit: a fast judgement suits screening many questions and flagging the uncertain ones in research pipelines. It does not replace deliberate clinical reasoning, and this model is not validated for any decision about a real patient (see the notice below).
Research use only — not a medical device
This model is released for research purposes only. It is not a medical device, has not been reviewed or cleared by any regulator, and has not been validated on any live clinical population. Do not use it for diagnosis, triage, treatment, trial eligibility, resource allocation or any other decision about a real person. It is distributed "as is", without warranty of any kind, express or implied, and without any representation of fitness for a particular purpose; the authors accept no liability for any use of it or of its outputs. Anyone evaluating it does so at their own risk and remains responsible for complying with the licences below and with all applicable law and research-ethics requirements. Its outputs are probability estimates from an incomplete picture of a patient, carry the limitations listed under Limitations, and must never substitute for the judgement of a qualified clinician.
Results — held-out test split (patient-disjoint)
Every row is the same test items under an identical protocol for all systems: same prompts, same option lists, same split. base is the untouched google/gemma-4-12B-it scored by its next-token distribution over the option numbers (it placed ≥ 99 % of its mass on valid options). Jev-Omni, when present, is the published akhilaaa3/Jev-Omni checkpoint run through its own loader on the same items. majority is the majority-class rate. The best system in each row is in bold; bootstrap 95 % CIs (500 resamples) are given for the pooled figures. 6 sources measured here are not tabulated — a model given only the patient's sex and age answers them as well as this one, so per-source numbers there would not measure what they appear to measure; they remain in the totals and in the run artifacts.
By task family
15,629 held-out test questions.
| Task family | majority | CareJev | base | Jev-Omni |
|---|---|---|---|---|
| 12-lead ECG (rendered) (15) | 0.737 | 0.792 | 0.587 | 0.529 |
| Heart sounds (3) | 0.694 | 0.700 | 0.429 | 0.404 |
| Chest X-ray | 0.490 | 0.870 | 0.694 | 0.571 |
| Sleep staging (rendered PSG) | 0.212 | 0.633 | 0.228 | 0.246 |
| ICU outcomes (SUPPORT2) (4) | 0.573 | 0.688 | 0.394 | 0.380 |
| EHR demo (eICU / MIMIC-IV) (5) | 0.465 | 0.460 | 0.279 | 0.242 |
| Exam questions (MedMCQA) | 0.303 | 0.728 | 0.715 | 0.708 |
| Transcription specialty | 0.264 | 0.445 | 0.417 | 0.421 |
| all test items (pooled) | 0.741 | 0.529 | 0.488 | |
| mean over sources | 0.739 | 0.561 | 0.528 |
CareJev 95 % CI: pooled 0.741 (0.734–0.747); mean over sources 0.739 (0.728–0.750).
Ranking quality (AUROC, yes/no questions, positive = "Yes")
| Task family | CareJev | base | Jev-Omni |
|---|---|---|---|
| 12-lead ECG (rendered) (15) | 0.812 | 0.672 | 0.656 |
| Heart sounds (3) | 0.626 | 0.495 | 0.543 |
| ICU outcomes (SUPPORT2) (4) | 0.829 | 0.650 | 0.651 |
| EHR demo (eICU / MIMIC-IV) (5) | 0.650 | 0.571 | 0.508 |
Calibration (expected calibration error, 10 equal-width bins; lower is better)
| Task family | CareJev | base | Jev-Omni |
|---|---|---|---|
| 12-lead ECG (rendered) (15) | 0.062 | 0.331 | 0.197 |
| Heart sounds (3) | 0.148 | 0.472 | 0.250 |
| Chest X-ray | 0.044 | 0.268 | 0.187 |
| Sleep staging (rendered PSG) | 0.066 | 0.491 | 0.145 |
| ICU outcomes (SUPPORT2) (4) | 0.037 | 0.571 | 0.453 |
| EHR demo (eICU / MIMIC-IV) (5) | 0.095 | 0.607 | 0.408 |
| Exam questions (MedMCQA) | 0.037 | 0.241 | 0.132 |
| Transcription specialty | 0.040 | 0.559 | 0.499 |
| all test items (pooled) | 0.018 | 0.394 | 0.250 |
| mean over sources | 0.092 | 0.367 | 0.228 |
Family rows weight each question by its test items (so accuracy is the pooled family accuracy); AUROC is the mean over the family's yes/no questions; the number in brackets is how many questions a family has (the datasets are credited under Data). The checkpoint was selected on the dev split only; the test split was never used for selection or calibration.
One-shot external benchmarks (never trained on, run once)
Every set below was converted to the typed-decision format and scored once, after the internal test, for CareJev, the untouched base model and the published Jev-Omni (through its own loader) on identical items. Nothing was tuned on them. They answer two different questions: the knowledge sets (MedQA, MMLU, PubMedQA, DecisionBench) ask what the fine-tune kept or lost of the base model's general ability; the shift sets (VQA-RAD, MedMNIST, CinC-2016, PTB-XL) ask whether the clinical training transfers to other hospitals, devices and question wordings.
Accuracy
| benchmark | majority | CareJev | base | Jev-Omni |
|---|---|---|---|---|
| MedQA-USMLE (4-option) | 0.277 | 0.738 | 0.687 | 0.701 |
| MMLU medical (8 subjects) | 0.304 | 0.821 | 0.821 | 0.812 |
| PubMedQA (yes/no/maybe) | 0.552 | 0.746 | 0.732 | 0.738 |
| DecisionBench medium | 0.365 | 0.792 | 0.761 | 0.867 |
| DecisionBench hard | 0.355 | 0.590 | 0.567 | 0.689 |
| VQA-RAD closed (radiology) | 0.530 | 0.733 | 0.753 | 0.781 |
| MedMNIST Pneumonia (CXR) | 0.630 | 0.602 | 0.876 | 0.836 |
| MedMNIST Derma (7-class) | 0.672 | 0.456 | 0.314 | 0.152 |
| MedMNIST Blood (8-class) | 0.178 | 0.368 | 0.412 | 0.344 |
| MedMNIST OrganA (11-class CT) | 0.174 | 0.288 | 0.334 | 0.310 |
| PhysioNet/CinC 2016 heart sounds | 0.800 | 0.617 | 0.200 | 0.200 |
| PTB-XL ECG (5 tasks) | 0.755 | 0.785 | 0.318 | 0.299 |
| all items (pooled) | 0.678 | 0.577 | 0.566 | |
| mean over benchmarks | 0.667 | 0.503 | 0.495 |
DecisionBench's own scoring (state-macro accuracy · confidence − accuracy gap): CareJev medium 0.821 (-14 pp), hard 0.555 (+2 pp) · base medium 0.771 (+19 pp), hard 0.551 (+37 pp) · Jev-Omni medium 0.882 (+1 pp), hard 0.643 (+15 pp)
Calibration (ECE, lower is better)
| benchmark | CareJev | base | Jev-Omni |
|---|---|---|---|
| MedQA-USMLE (4-option) | 0.054 | 0.209 | 0.107 |
| MMLU medical (8 subjects) | 0.038 | 0.161 | 0.100 |
| PubMedQA (yes/no/maybe) | 0.029 | 0.244 | 0.171 |
| DecisionBench medium | 0.146 | 0.198 | 0.030 |
| DecisionBench hard | 0.118 | 0.380 | 0.165 |
| VQA-RAD closed (radiology) | 0.067 | 0.213 | 0.083 |
| MedMNIST Pneumonia (CXR) | 0.183 | 0.106 | 0.033 |
| MedMNIST Derma (7-class) | 0.048 | 0.580 | 0.549 |
| MedMNIST Blood (8-class) | 0.110 | 0.529 | 0.372 |
| MedMNIST OrganA (11-class CT) | 0.068 | 0.481 | 0.277 |
| PhysioNet/CinC 2016 heart sounds | 0.076 | 0.608 | 0.300 |
| PTB-XL ECG (5 tasks) | 0.082 | 0.635 | 0.496 |
| all items (pooled) | 0.014 | 0.354 | 0.228 |
| mean over benchmarks | 0.084 | 0.430 | 0.292 |
PTB-XL: compare AUROC, not accuracy. We render the PTB-XL ECGs in the 12-lead paper layout this model was trained on. That layout drags the base model's accuracy down (hypertrophy 0.176, against 0.856 on a plainer strip layout) without changing how well it ranks cases, so accuracy would overstate the margin. AUROC does not depend on a decision threshold, and CareJev has the highest AUROC on all 5 tasks, against either layout:
| PTB-XL task (AUROC) | CareJev | base | Jev-Omni | base, strip layout |
|---|---|---|---|---|
| conduction disturbance | 0.665 | 0.603 | 0.634 | 0.605 |
| hypertrophy | 0.664 | 0.442 | 0.447 | 0.603 |
| myocardial infarction | 0.781 | 0.663 | 0.653 | 0.670 |
| normal ECG | 0.844 | 0.594 | 0.598 | 0.678 |
| ST/T changes | 0.820 | 0.399 | 0.474 | 0.535 |
Caveats that belong next to those numbers: MedQA and MMLU are almost certainly in the base model's pre-training; MedMNIST-Pneumonia draws on the same paediatric CXR collection as the "viral pneumonia" class of our COVID-19 Radiography training data, so its number may be optimistic; and on CinC-2016 heart sounds no model beats the trivial baseline: answering "normal" to everything scores 0.800, this model scores 0.617 with an AUROC of 0.542, and the base model and Jev-Omni score 0.200 because they answer "abnormal" to almost every recording. Being better than the other two there is not the same as being useful: the honest reading is that the heart-sound pathway does not transfer to this collection, and it is the open item for the next run.
Further analyses
Post-hoc calibration
Fitted on the dev split: a global temperature (T = 1.036), with per-source temperature / yes-bias terms where they helped. Not shipped. The fit was validated on the single-prompt path (test NLL 0.6377 to 0.6284), but the released path averages over option rotations, which already smooths the probabilities: it measures an expected calibration error of 0.018 without any post-hoc term. Shipping a correction that was never validated on the path we release would make the tables describe a configuration nobody runs, so the loader returns the raw head probabilities and every number on this card is measured that way. The fit is kept with the run artifacts for the record.
| before | after (fitted, not shipped) | |
|---|---|---|
| dev NLL (fit set) | 0.6740 | 0.6738 |
| test NLL | 0.6377 | 0.6284 |
| test ECE | 0.0218 | 0.0096 |
| test mean-source NLL | 0.6127 | 0.6084 |
| test accuracy | 0.7375 | 0.7389 |
Selective prediction (does confidence rank errors?)
Accuracy when only the most confident share of test questions is answered; a model whose confidence is informative gains accuracy as coverage falls.
15,629 test questions (text 5,847, image 8,835, audio 947).
| group | acc @ 50 % | acc @ 80 % | acc @ 100 % |
|---|---|---|---|
| all | 0.897 | 0.815 | 0.741 |
| text | 0.868 | 0.752 | 0.670 |
| image | 0.921 | 0.846 | 0.791 |
| audio | 0.829 | 0.748 | 0.711 |
Beyond accuracy (balanced accuracy and score questions)
On skewed questions accuracy can sit at the majority rate while the model still separates the
classes; balanced accuracy (the mean of per-class recall) shows that. For score questions the
interface returns the expected level E[level] = Σ i·p_i.
| Task family | CareJev | base | Jev-Omni |
|---|---|---|---|
| 12-lead ECG (rendered) (15) | 0.630 | 0.570 | 0.546 |
| Heart sounds (3) | 0.515 | 0.501 | 0.505 |
| Chest X-ray | 0.798 | 0.457 | 0.434 |
| Sleep staging (rendered PSG) | 0.651 | 0.225 | 0.255 |
| ICU outcomes (SUPPORT2) (4) | 0.590 | 0.434 | 0.423 |
| EHR demo (eICU / MIMIC-IV) (5) | 0.222 | 0.146 | 0.130 |
| Exam questions (MedMCQA) | 0.725 | 0.710 | 0.695 |
| Transcription specialty | 0.373 | 0.555 | 0.514 |
Score questions, mean absolute error of the expected level: eicu_demo_los 1.40 levels (majority 1.77); mimic4_demo_los 2.31 levels (majority 2.38); support2_los 1.53 levels (majority 1.57).
Sub-groups (SUPPORT2, as recorded in the state)
2,739 SUPPORT2 test questions (in-hospital death, 2- and 6-month survival); groups with fewer than 30 questions are not shown.
| group | accuracy | AUROC | ECE |
|---|---|---|---|
| sex: female | 0.758 | 0.841 | 0.019 |
| sex: male | 0.790 | 0.869 | 0.030 |
| age: <50 | 0.803 | 0.893 | 0.043 |
| age: 50-64 | 0.791 | 0.862 | 0.034 |
| age: 65-79 | 0.747 | 0.828 | 0.014 |
| age: 80+ | 0.780 | 0.863 | 0.057 |
| race: black | 0.803 | 0.873 | 0.036 |
| race: hispanic | 0.758 | 0.892 | 0.138 |
| race: other | 0.833 | 0.863 | 0.108 |
| race: white | 0.769 | 0.850 | 0.025 |
Latency of the released path
| path | questions | ms / question | input tokens / s |
|---|---|---|---|
| noul (Yes/No), one prompt | 40 | 307 | 925 |
| score, 10 levels, released path (8 rotations, one batch) | 40 | 2287 | 154 |
| score, 10 levels, single prompt (permutations=0) | 40 | 338 | 1042 |
| prefix cache (opt-in), 3-4 questions per state, single prompt each | 40 | 183 | - |
Measured on NVIDIA GB10.
Data
CareJev-Omni was trained only on publicly available, open-access medical datasets: rendered 12-lead ECGs, chest X-rays, heart-sound recordings, sleep recordings, ICU outcome records, de-identified EHR demo data, medical exam questions and clinical transcriptions. No credentialed data and no protected health information were used, and evaluation uses a patient-disjoint held-out split.
Attribution: PhysioNet/CinC Challenge 2020, CirCor DigiScope, Sleep-EDF, eICU-CRD demo and MIMIC-IV demo (PhysioNet); COVID-19 Radiography Database (CC BY 4.0); BMD-HS (CC BY 4.0); SUPPORT2 (Vanderbilt / UCI); MedMCQA (MIT); Medical Transcriptions (CC0). Each dataset's own licence and terms continue to apply.
Quick start
pip install -r requirements.txt # torch, torchvision, transformers==5.17.0, peft, ...
# audio input also needs ffmpeg on PATH
torchvision is required even for text-only use: transformers' Gemma 4 image processor imports it at
module load. This repo is self-contained: the frozen base weights, the adapter applied on top at load and the decision head. No network access is needed after download. The low-rank delta is deliberately not folded into the weights: merging it into bf16 leaves accuracy and NLL intact but costs about half of this model's calibration and flips 3.4 % of answers (measured, 300 test items), and calibration is the property this card is about. python example.py (shipped in this repo) prints one answer of each question type.
Every file's SHA-256 is in sha256.json; to check a download:
python -c "import json,hashlib,pathlib;[print(k, hashlib.sha256(pathlib.Path(k).read_bytes()).hexdigest()==v) for k,v in json.load(open('sha256.json')).items()]"
from huggingface_hub import snapshot_download
import sys; sys.path.insert(0, snapshot_download("nevermindai/carejev-omni", allow_patterns=["*.py"]))
from carejev_omni_loader import load_carejev_omni # same contract as Jev-Omni's load_jev_omni
clf = load_carejev_omni("nevermindai/carejev-omni")
print(clf.predict(
state="71-year-old male. Admission type: Emergency. ... DAY-3 PHYSIOLOGY: mean arterial pressure 55 mmHg; ...",
question="Will this patient die before being discharged from this hospitalization?",
options=["Yes", "No"],
))
# {'prediction': 'Yes', 'prediction_index': 0, 'confidence': 0.71, 'probabilities': {'Yes': 0.71, 'No': 0.29}}
The inference code ships in this repo (carejev_omni_loader.py and the carejev_omni/ package it
imports), so pip install -r requirements.txt is all a fresh environment needs.
How to use
Once loaded, clf.predict(...) answers one question about one patient state, and
clf.predict_many(...) answers several questions about the same state.
Inputs
| argument | type | meaning |
|---|---|---|
state |
str |
what is known about the patient, written as text (notes, values, context) |
question |
str |
the question to answer |
options |
list[str] |
the allowed answers: 2 to 256, all different |
media |
path, optional | an image (chest X-ray, rendered 12-lead ECG, spectrogram) or a heart-sound recording (.wav) |
modality |
"text" · "image" · "audio" |
what media is; default "text" (no media) |
qtype |
"noul" · "choice" · "score", optional |
inferred when omitted (["Yes", "No"] is a yes/no question, anything else is a choice); pass "score" when the options are ordered levels |
permutations |
None or 0 |
None (default) averages over rotations of the options, so their order cannot change the answer; 0 scores one prompt in your order (faster) |
Output: a plain, JSON-serialisable dict:
{
"prediction": "Yes",
"prediction_index": 0,
"confidence": 0.71,
"probabilities": {"Yes": 0.71, "No": 0.29}
}
probabilities holds every option, in the order you passed them, summing to 1. prediction is the
most probable option, prediction_index its position in your list and confidence its probability.
A threshold on confidence is how you decide which answers to use as they are and which to pass on
for review.
The same structure comes back for every question type (values below are illustrative). Choice:
{
"prediction": "Sepsis",
"prediction_index": 0,
"confidence": 0.82,
"probabilities": {"Sepsis": 0.82, "Heart failure": 0.09, "Pulmonary embolism": 0.06,
"Diabetic ketoacidosis": 0.03}
}
Score (ordered levels): the most probable level, plus every level's probability:
{
"prediction": "8 to 14 days",
"prediction_index": 8,
"confidence": 0.34,
"probabilities": {"Less than 1 day": 0.0, "1 day": 0.01, "2 days": 0.02, "3 days": 0.04,
"4 days": 0.06, "5 days": 0.08, "6 days": 0.09, "7 days": 0.11,
"8 to 14 days": 0.34, "More than 14 days": 0.25}
}
predict_many: a list of the same dicts, one per question, in question order:
[
{"prediction": "No", "prediction_index": 1, "confidence": 0.77, "probabilities": {"Yes": 0.23, "No": 0.77}},
{"prediction": "Yes", "prediction_index": 0, "confidence": 0.58, "probabilities": {"Yes": 0.58, "No": 0.42}}
]
Examples
state = "PATIENT: 68-year-old male. ... DAY-3 PHYSIOLOGY: mean arterial pressure 72 mmHg; ..."
# yes / no: read the probability of "Yes"
out = clf.predict(state=state, options=["Yes", "No"],
question="Will this patient die before being discharged from this hospitalization?")
p_yes = out["probabilities"]["Yes"]
# choice: one of several options
out = clf.predict(state=state, question="Which single problem best explains this presentation?",
options=["Sepsis", "Heart failure", "Pulmonary embolism", "Diabetic ketoacidosis"])
answer, confidence = out["prediction"], out["confidence"]
# score: ordered levels; the expected level is the probability-weighted position
levels = ["Less than 1 day", "1 day", "2 days", "3 days", "4 days", "5 days", "6 days", "7 days",
"8 to 14 days", "More than 14 days"]
out = clf.predict(state=state, question="How long will this patient's hospital stay be in total?",
options=levels, qtype="score")
expected_level = sum(i * out["probabilities"][o] for i, o in enumerate(levels))
# an image (here a chest X-ray) or a heart-sound recording (audio needs ffmpeg on PATH)
clf.predict(state="Frontal chest radiograph.", question="What does this chest X-ray most likely show?",
options=["COVID-19", "Normal", "Lung opacity", "Viral pneumonia"],
media="cxr.png", modality="image")
clf.predict(state="Phonocardiogram: digital stethoscope recording of heart sounds.",
question="Is a heart murmur audible in this recording?", options=["Yes", "No"],
media="pcg.wav", modality="audio")
# several questions about the same state: a list of the same dicts, in question order
outs = clf.predict_many(state=state, questions=[
("Will this patient die before being discharged from this hospitalization?", ["Yes", "No"]),
("Will this patient be alive 6 months (180 days) after study entry?", ["Yes", "No"]),
])
Questions phrased like the training data work best: the examples above use that wording, and the
model card's task families show what it was trained on. predict_many(..., use_prefix_cache=True)
encodes the state once for all the questions; it is faster but scores each question with a single
prompt in your option order, like permutations=0. Every number on this card uses the default path.
Limitations
- Trained on open, mostly small datasets; SUPPORT2 (1989–94, five US hospitals) dominates the text portion. Not validated on any live clinical population.
- Signal "images" are our own renderings; performance depends on that rendering.
- Heart-sound questions whose label age and sex alone predict as well as this model (
bmdhs_ar,bmdhs_as,bmdhs_mr,bmdhs_normal) are not reported: a per-source number there would not measure the recording. The heart-sound rows that are tabulated (bmdhs_ms,circor_murmur,circor_outcome) do not have that shortcut, and on the external CinC-2016 collection no model beats the trivial answer. - Video is supported by the architecture but received no training data.
- eICU / MIMIC-IV demo sources are tiny (≤ 100 patients); their numbers are noisy.
- Calibration is reported on the internal test split only; distribution shift will degrade it.
- For a single structured-tabular task, a tabular model is at least as good. On the SUPPORT2 outcomes, a gradient-boosted model fitted on the raw numeric columns available at study entry is level with this model on the three binary outcomes (AUROC within 0.01, accuracy within 0.01) and clearly better on the length-of-stay score question, with slightly lower NLL on all four. Per-source classical models also win on a few other sources (notably the tiny eICU demo tasks and CirCor murmur). This model is for the case where the evidence is mixed — prose, rendered signals, sounds — and one interface has to cover all of it, not for replacing a well-fitted tabular risk score.
Relation to other work
Independent open model implementing the typed-decision interface described by TypeSafe AI's Jev; not affiliated with or trained on outputs of TypeSafe AI, Jev, Jev-Omni, or Google. Architecture and prompt follow Jev-Omni for drop-in compatibility.
Evaluation protocol
- Splits are patient-disjoint and were fixed before any model saw the data.
- The test split is touched once per model, after training; selection uses dev only.
- Baselines share the exact prompt text and option lists (the base model's answer is read from its next-token distribution over the option numbers).
- Published comparisons are to the untouched base model and to the other released model implementing this interface. Non-neural baselines (TF-IDF / tabular / pixel / MFCC models fitted per source) were measured and are kept with the run artifacts, but are not part of the published tables: they are bespoke per task and per modality, so they do not answer the question this card is about, which is what one general model does across every source. See Limitations for where such a model is the better tool.
- Intervals: non-parametric bootstrap over test examples, 500 resamples, percentile 95 % CI.
- Calibration terms are fitted on dev and reported on test before/after.
- Every number on this card is rendered directly from the evaluation outputs of this checkpoint.
Licence
The base weights are Gemma 4, published by Google under Apache-2.0 (licence). Our own contributions — decision head, loader code and this card — are Apache-2.0 as well, so the whole release is Apache-2.0. Dataset licences are listed above and remain separate. Both licences disclaim warranties and limit liability; nothing here grants any additional warranty or indemnity, and the research-use notice at the top of this card applies to every part of the release. Card generated 2026-10-04 from run artifacts.
- Downloads last month
- 72
