Med Advisor Conversation 4B

A 4B medical explainer that adjusts its answer to who is asking: patient, caregiver, science-literate adult, medical student or healthcare worker.

It is Qwen3-4B, fully fine-tuned on medical conversations. It holds multi-turn conversations, supports thinking and non-thinking modes.

Scores for the base model, SFT and RL100 in both modes, graded by GPT-6 Luna

HealthBench Consensus, 1,000 English prompts Qwen3-4B (base) After fine-tuning (SFT) Med Advisor (RL100)
Non-thinking 34.5% 40.7% 54.4%
Thinking 32.0% 47.9% 59.9%

RL100 is the released model: the reinforcement-learning (RL) training checkpoint after 100 updates. Same system message and sampling settings for every model. Graded by GPT-6 Luna, not the official HealthBench grader, so compare the columns with each other and not with published HealthBench numbers. Details.

This model has not been clinically validated. It is for education and research. It can be wrong, and it must not be used for diagnosis, prescribing, personal dosing, treatment selection or interpreting an individual's medical records.

Three examples · Scope · Quick start · All examples · RL Recipe · Evaluation · Limitations

Start with three conversations

Scope

Audiences (5) Curious patient · Caregiver · Science-literate adult · Medical student · Healthcare worker
Kinds of request (6) Explanation, mechanism and evidence · Lab, test and risk interpretation · Medication and numerical reasoning · Self-care and next steps · Safety escalation · Misinformation correction
Domains (16) General medicine · Oncology · Laboratory and diagnostic medicine · Medication safety · Cardiovascular · Infectious disease · Endocrinology · Neurology · Respiratory · Musculoskeletal · Gastroenterology · Dermatology · Genetics and genomics · Nephrology and urology · Rare disease · Mental health
Modes Thinking (reasons first, then answers) and non-thinking, switched with enable_thinking
Language English

The audience changes the depth and vocabulary of an answer. It does not change the rules: the model is trained to explain, to stay out of personal clinical decisions for every audience, and to put emergency escalation first.

Quick start

pip install "transformers>=4.57" accelerate safetensors

The BF16 weights need about 8 GB of memory, plus room for the context.

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

MODEL_ID = "vmal/med-advisor-conversation-4B"
THINKING = False  # True: the model reasons before the visible answer

tokenizer = AutoTokenizer.from_pretrained(MODEL_ID)
model = AutoModelForCausalLM.from_pretrained(MODEL_ID, dtype=torch.bfloat16, device_map="auto").eval()

system_message = (
    "You are Med Advisor, a clear, evidence-aware medical and scientific "
    "explainer. Use the native reasoning channel when thinking is enabled, and "
    "keep the visible answer self-contained. Give useful general education with "
    "calibrated uncertainty. Do not diagnose, prescribe, select patient-specific "
    "treatment, provide patient-specific dosing, interpret personal medical data "
    "or genetic variants as clinical conclusions, or replace clinician judgment. "
    "Escalate emergencies immediately."
)
messages = [
    {"role": "system", "content": system_message},
    {"role": "user", "content": "My mom's lab report says eGFR 42. What does that number mean?"},
]
inputs = tokenizer.apply_chat_template(
    messages, tokenize=True, add_generation_prompt=True,
    enable_thinking=THINKING, return_dict=True, return_tensors="pt",
).to(model.device)

with torch.inference_mode():
    output = model.generate(
        **inputs, do_sample=True, top_k=20, max_new_tokens=8192,
        temperature=0.6 if THINKING else 0.7,
        top_p=0.95 if THINKING else 0.8,
    )

text = tokenizer.decode(output[0, inputs["input_ids"].shape[-1]:], skip_special_tokens=True)
answer = text.split("</think>")[-1].strip()
print(answer)
messages.append({"role": "assistant", "content": answer})  # keep only the visible answer in history

Three things matter for getting the behavior shown here:

  • Use this system message. The model was trained and evaluated with it.
  • Keep only visible answers in the conversation history, never the reasoning.
  • Cap max_new_tokens in thinking mode. The model occasionally loops while reasoning. If an answer does not finish, retry in non-thinking mode.

Example conversations

Twenty selected conversations show how the model answers five audiences across six kinds of request, in thinking and non-thinking mode. Browse all examples.

Illustrative examples, selected to show response style. The model's replies are unedited. Not medical advice.

One short excerpt. A caregiver has just been given seizure first-aid steps and asks for something shorter:

User: That makes sense, thank you — so no holding her down, nothing in her mouth, and time it. Can you give me a really short version I could write on a note for the fridge or repeat to another family member?

Med Advisor:

Fridge note:

  • Don’t hold her down.
  • Don’t put anything in her mouth.
  • Clear hard objects.
  • If on the floor, turn her on her side if safe.
  • Time it.
  • Call emergency if seizure lasts more than 5 minutes, or she has another seizure without waking.
  • Follow her seizure action plan.

RL Recipe

Evaluation

HealthBench Consensus scores for the base model, SFT and the RL training checkpoints

Reading the chart. Baseline is Qwen3-4B before any training. SFT is the model after supervised fine-tuning. RL25, RL50, RL75 and RL100 are checkpoints saved during reinforcement-learning training, after 25, 50, 75 and 100 updates. RL100 is the released model. T is thinking mode and NT is non-thinking mode.

What was run. 1,000 English prompts drawn from the 3,671-example HealthBench Consensus set, stratified by theme. Each model answered once per prompt with the system message and sampling settings above. Only the visible answer was graded. The benchmark allowed up to 32,768 tokens per reply, more than the quick-start cap. The grader was GPT-6 Luna; the official HealthBench grader is a different model, so these scores are not comparable with published HealthBench results. Full protocol, intervals and saved results.

Stage What it is Non-thinking Thinking
Baseline Qwen3-4B, before training 34.5% 32.0%
SFT After supervised fine-tuning 40.7% 47.9%
RL25 RL training checkpoint, 25 updates 41.9% 51.8%
RL50 RL training checkpoint, 50 updates 48.9% 55.6%
RL75 RL training checkpoint, 75 updates 51.5% 56.9%
RL100 RL training checkpoint, 100 updates (this model) 54.4% 59.9%

The score rises at every RL checkpoint in both modes. The 95% intervals are about ±2.5 points; for RL100 they are 51.8 to 56.8 (non-thinking) and 57.4 to 62.4 (thinking).

Against the base model in the same mode, RL100 gains 19.9 points without thinking (paired 95% interval 17.4 to 22.3) and 28.0 points with thinking (25.5 to 30.5).

By theme (mean score per prompt):

Theme Base, non-thinking Med Advisor, non-thinking Med Advisor, thinking
Emergency referrals 22.0% 55.3% 65.9%
Context seeking 22.5% 50.5% 58.1%
Communication 10.8% 32.8% 42.6%
Hedging 51.2% 70.1% 77.0%
Global health 62.4% 79.8% 82.4%
Health data tasks 41.7% 50.9% 49.1%
Complex responses 21.0% 27.3% 25.6%

Reading these numbers. Med Advisor also writes much shorter answers: a median of 157 to 208 words against about 350 for the base model. The gain is not only brevity. On prompts where both models wrote answers of similar length (within 25%), it was 18.3 points without thinking (194 prompts) and 16.1 points with thinking (64 prompts).

Intended use and limitations

Use it for learning and explaining: what a term or result means, how a mechanism works, how strong the evidence is, what to ask a clinician, and when something is an emergency. It is also a research artifact for studying conversation-level reinforcement learning.

Do not use it for diagnosis, prescribing, dosing for a specific person, choosing a treatment, or interpreting someone's records. It is trained to decline these, but it will not always do so.

License

Apache 2.0, the same license as the base model.

Citation

If you use this checkpoint, please cite the model page.

Downloads last month
304
Safetensors
Model size
4B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for vmal/med-advisor-conversation-4B

Finetuned
Qwen/Qwen3-4B
Finetuned
(1133)
this model
Quantizations
2 models