Instructions to use vmal/med-advisor-conversation-4B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use vmal/med-advisor-conversation-4B with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="vmal/med-advisor-conversation-4B") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("vmal/med-advisor-conversation-4B") model = AutoModelForCausalLM.from_pretrained("vmal/med-advisor-conversation-4B", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use vmal/med-advisor-conversation-4B with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "vmal/med-advisor-conversation-4B" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "vmal/med-advisor-conversation-4B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/vmal/med-advisor-conversation-4B
- SGLang
How to use vmal/med-advisor-conversation-4B with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "vmal/med-advisor-conversation-4B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "vmal/med-advisor-conversation-4B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "vmal/med-advisor-conversation-4B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "vmal/med-advisor-conversation-4B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use vmal/med-advisor-conversation-4B with Docker Model Runner:
docker model run hf.co/vmal/med-advisor-conversation-4B
Med Advisor Conversation 4B
A 4B medical explainer that adjusts its answer to who is asking: patient, caregiver, science-literate adult, medical student or healthcare worker.
It is Qwen3-4B, fully fine-tuned on medical conversations. It holds multi-turn conversations, supports thinking and non-thinking modes.
| HealthBench Consensus, 1,000 English prompts | Qwen3-4B (base) | After fine-tuning (SFT) | Med Advisor (RL100) |
|---|---|---|---|
| Non-thinking | 34.5% | 40.7% | 54.4% |
| Thinking | 32.0% | 47.9% | 59.9% |
RL100 is the released model: the reinforcement-learning (RL) training checkpoint after 100 updates. Same system message and sampling settings for every model. Graded by GPT-6 Luna, not the official HealthBench grader, so compare the columns with each other and not with published HealthBench numbers. Details.
This model has not been clinically validated. It is for education and research. It can be wrong, and it must not be used for diagnosis, prescribing, personal dosing, treatment selection or interpreting an individual's medical records.
Three examples · Scope · Quick start · All examples · RL Recipe · Evaluation · Limitations
Start with three conversations
- A number explained in plain terms: The advert says 50% fewer heart attacks. The model turns a 50% relative reduction at a 4% baseline risk into absolute numbers, then writes a version to say to a patient.
- A correction changes the answer: Metformin or metoprolol? The user checks the bottle in turn three, and the model updates its explanation.
- A follow-up changes the urgency: An eGFR of 42 on a weekend, then new confusion. When the user adds acute symptoms, the next reply leads with escalation.
Scope
| Audiences (5) | Curious patient · Caregiver · Science-literate adult · Medical student · Healthcare worker |
| Kinds of request (6) | Explanation, mechanism and evidence · Lab, test and risk interpretation · Medication and numerical reasoning · Self-care and next steps · Safety escalation · Misinformation correction |
| Domains (16) | General medicine · Oncology · Laboratory and diagnostic medicine · Medication safety · Cardiovascular · Infectious disease · Endocrinology · Neurology · Respiratory · Musculoskeletal · Gastroenterology · Dermatology · Genetics and genomics · Nephrology and urology · Rare disease · Mental health |
| Modes | Thinking (reasons first, then answers) and non-thinking, switched with enable_thinking |
| Language | English |
The audience changes the depth and vocabulary of an answer. It does not change the rules: the model is trained to explain, to stay out of personal clinical decisions for every audience, and to put emergency escalation first.
Quick start
pip install "transformers>=4.57" accelerate safetensors
The BF16 weights need about 8 GB of memory, plus room for the context.
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
MODEL_ID = "vmal/med-advisor-conversation-4B"
THINKING = False # True: the model reasons before the visible answer
tokenizer = AutoTokenizer.from_pretrained(MODEL_ID)
model = AutoModelForCausalLM.from_pretrained(MODEL_ID, dtype=torch.bfloat16, device_map="auto").eval()
system_message = (
"You are Med Advisor, a clear, evidence-aware medical and scientific "
"explainer. Use the native reasoning channel when thinking is enabled, and "
"keep the visible answer self-contained. Give useful general education with "
"calibrated uncertainty. Do not diagnose, prescribe, select patient-specific "
"treatment, provide patient-specific dosing, interpret personal medical data "
"or genetic variants as clinical conclusions, or replace clinician judgment. "
"Escalate emergencies immediately."
)
messages = [
{"role": "system", "content": system_message},
{"role": "user", "content": "My mom's lab report says eGFR 42. What does that number mean?"},
]
inputs = tokenizer.apply_chat_template(
messages, tokenize=True, add_generation_prompt=True,
enable_thinking=THINKING, return_dict=True, return_tensors="pt",
).to(model.device)
with torch.inference_mode():
output = model.generate(
**inputs, do_sample=True, top_k=20, max_new_tokens=8192,
temperature=0.6 if THINKING else 0.7,
top_p=0.95 if THINKING else 0.8,
)
text = tokenizer.decode(output[0, inputs["input_ids"].shape[-1]:], skip_special_tokens=True)
answer = text.split("</think>")[-1].strip()
print(answer)
messages.append({"role": "assistant", "content": answer}) # keep only the visible answer in history
Three things matter for getting the behavior shown here:
- Use this system message. The model was trained and evaluated with it.
- Keep only visible answers in the conversation history, never the reasoning.
- Cap
max_new_tokensin thinking mode. The model occasionally loops while reasoning. If an answer does not finish, retry in non-thinking mode.
Example conversations
Twenty selected conversations show how the model answers five audiences across six kinds of request, in thinking and non-thinking mode. Browse all examples.
Illustrative examples, selected to show response style. The model's replies are unedited. Not medical advice.
One short excerpt. A caregiver has just been given seizure first-aid steps and asks for something shorter:
User: That makes sense, thank you — so no holding her down, nothing in her mouth, and time it. Can you give me a really short version I could write on a note for the fridge or repeat to another family member?
Med Advisor:
Fridge note:
- Don’t hold her down.
- Don’t put anything in her mouth.
- Clear hard objects.
- If on the floor, turn her on her side if safe.
- Time it.
- Call emergency if seizure lasts more than 5 minutes, or she has another seizure without waking.
- Follow her seizure action plan.
RL Recipe
- RL training recipe and code: architecture, simulated conversations, grading rubrics, reward calculation, and policy updates.
Evaluation
Reading the chart. Baseline is Qwen3-4B before any training. SFT is the model after supervised fine-tuning. RL25, RL50, RL75 and RL100 are checkpoints saved during reinforcement-learning training, after 25, 50, 75 and 100 updates. RL100 is the released model. T is thinking mode and NT is non-thinking mode.
What was run. 1,000 English prompts drawn from the 3,671-example HealthBench Consensus set, stratified by theme. Each model answered once per prompt with the system message and sampling settings above. Only the visible answer was graded. The benchmark allowed up to 32,768 tokens per reply, more than the quick-start cap. The grader was GPT-6 Luna; the official HealthBench grader is a different model, so these scores are not comparable with published HealthBench results. Full protocol, intervals and saved results.
| Stage | What it is | Non-thinking | Thinking |
|---|---|---|---|
| Baseline | Qwen3-4B, before training | 34.5% | 32.0% |
| SFT | After supervised fine-tuning | 40.7% | 47.9% |
| RL25 | RL training checkpoint, 25 updates | 41.9% | 51.8% |
| RL50 | RL training checkpoint, 50 updates | 48.9% | 55.6% |
| RL75 | RL training checkpoint, 75 updates | 51.5% | 56.9% |
| RL100 | RL training checkpoint, 100 updates (this model) | 54.4% | 59.9% |
The score rises at every RL checkpoint in both modes. The 95% intervals are about ±2.5 points; for RL100 they are 51.8 to 56.8 (non-thinking) and 57.4 to 62.4 (thinking).
Against the base model in the same mode, RL100 gains 19.9 points without thinking (paired 95% interval 17.4 to 22.3) and 28.0 points with thinking (25.5 to 30.5).
By theme (mean score per prompt):
| Theme | Base, non-thinking | Med Advisor, non-thinking | Med Advisor, thinking |
|---|---|---|---|
| Emergency referrals | 22.0% | 55.3% | 65.9% |
| Context seeking | 22.5% | 50.5% | 58.1% |
| Communication | 10.8% | 32.8% | 42.6% |
| Hedging | 51.2% | 70.1% | 77.0% |
| Global health | 62.4% | 79.8% | 82.4% |
| Health data tasks | 41.7% | 50.9% | 49.1% |
| Complex responses | 21.0% | 27.3% | 25.6% |
Reading these numbers. Med Advisor also writes much shorter answers: a median of 157 to 208 words against about 350 for the base model. The gain is not only brevity. On prompts where both models wrote answers of similar length (within 25%), it was 18.3 points without thinking (194 prompts) and 16.1 points with thinking (64 prompts).
Intended use and limitations
Use it for learning and explaining: what a term or result means, how a mechanism works, how strong the evidence is, what to ask a clinician, and when something is an emergency. It is also a research artifact for studying conversation-level reinforcement learning.
Do not use it for diagnosis, prescribing, dosing for a specific person, choosing a treatment, or interpreting someone's records. It is trained to decline these, but it will not always do so.
License
Apache 2.0, the same license as the base model.
Citation
If you use this checkpoint, please cite the model page.
- Downloads last month
- 304

