Instructions to use tbuckley/PrecepTron-32B-CPC with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use tbuckley/PrecepTron-32B-CPC with PEFT:
from peft import PeftModel from transformers import AutoModelForCausalLM base_model = AutoModelForCausalLM.from_pretrained("Qwen/Qwen3-32B") model = PeftModel.from_pretrained(base_model, "tbuckley/PrecepTron-32B-CPC") - Notebooks
- Google Colab
- Kaggle
PrecepTron-32B-CPC
PrecepTron-32B-CPC is a fine-tuned version of Qwen/Qwen3-32B
that acts as an automated rubric-based grader ("LLM-as-a-judge") for clinical
reasoning. It is fine-tuned via LoRA and distributed here as a PEFT adapter that
loads on top of the Qwen/Qwen3-32B base weights. It is one of four task-specific
PrecepTron-32B judges released alongside the paper Scaling Clinical Judgment to Evaluate Medical AI.
Given a clinical case and a response to evaluate, the model outputs a numeric score on the task's rubric together with a short justification, as a JSON object.
- Task: NEJM Clinicopathological Conferences (CPCs) — differential-diagnosis quality on the 0–5 Bond score. The same adapter is also used to grade the BIDMC emergency-department triage task.
- Base model: Qwen/Qwen3-32B
- Method: LoRA supervised fine-tuning (PEFT) on physician rubric scores
- LoRA config: r=16, α=32, dropout=0.05, 3 epochs, bin-stratified ("balanced") oversampling
- Output format:
{"score": <number>, "justification": "<text>"} - Website: https://preceptron.net · Code: https://github.com/2v/PrecepTron · Benchmark: https://huggingface.co/datasets/tbuckley/GRAND-ROUNDS
Usage
import torch
from peft import PeftModel
from transformers import AutoModelForCausalLM, AutoTokenizer
base = AutoModelForCausalLM.from_pretrained(
"Qwen/Qwen3-32B", torch_dtype=torch.bfloat16, device_map="auto")
model = PeftModel.from_pretrained(base, "tbuckley/PrecepTron-32B-CPC")
tok = AutoTokenizer.from_pretrained("Qwen/Qwen3-32B")
messages = [
{"role": "system", "content": SYSTEM_PROMPT}, # see "Prompt" below
{"role": "user", "content": USER_PROMPT}, # case + response to score
]
inputs = tok.apply_chat_template(
messages, add_generation_prompt=True, return_tensors="pt").to(model.device)
out = model.generate(inputs, max_new_tokens=512, do_sample=False)
print(tok.decode(out[0][inputs.shape[1]:], skip_special_tokens=True))
Prompt
The model was trained with the exact system and user prompts below — the same prompts it expects at inference time. The scoring rubric is fixed and baked into the system prompt below.
System prompt
You are an expert clinical evaluator. You will be given a clinical response and a scoring rubric.
Score the response according to the rubric. You must output ONLY valid JSON with your scoring result.
Your output must be a JSON object with these fields:
- "score": the numeric score you assign
- "justification": a brief explanation of your scoring decision
## Bond Score Rubric
The scale does not contain a score of 1.
- 0: no suggestions close to the target diagnosis
- 2: the suggestions included something related, but unlikely to be helpful
- 3: the suggestions included something closely related that might have been helpful
- 4: the suggestions included something very close, but not exact
- 5: the actual diagnosis was suggested in the differential
User prompt template (case-specific fields filled in at inference)
## Final Diagnosis
{final_diagnosis}
## Response to Score
{response}
Training data
Supervised fine-tuning on physician rubric scores from the study dataset. The assistant target is the reconciled physician score (or, when no reconciliation exists, a single physician's score selected deterministically), validated against the task's allowed score scale. Train/validation case IDs were held out of all reported evaluations.
Limitations
This is a research artifact for the automated evaluation of medical AI. It is not a diagnostic or clinical decision-support tool and must not be used for patient care. Scores reflect agreement with the specific rubrics and physician annotations in this study and may not transfer to other rubrics, populations, or response formats.
Citation
@article{buckley2026preceptron,
title = {Scaling Clinical Judgment to Evaluate Medical AI},
author = {Buckley, Thomas A. and Kanjee, Zahir and Brodeur, Peter G. and
Crowe, Byron and Pettinato, Anthony M. and Shah, Aashna P. and
Haimovich, Adrian D. and McCoy, Liam G. and Restrepo, Daniel and
Goh, Ethan and Chen, Jonathan H. and Zwaan, Laura and
Goodman, Katherine E. and Morgan, Daniel J. and
Abdulnour, Raja-Elie E. and Rodman, Adam and Manrai, Arjun K.},
year = {2026},
note = {Preprint, forthcoming}
}
- Downloads last month
- 2
Model tree for tbuckley/PrecepTron-32B-CPC
Base model
Qwen/Qwen3-32B