Instructions to use ZeroOneAI/ZEO-Med-2 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use ZeroOneAI/ZEO-Med-2 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="ZeroOneAI/ZEO-Med-2")# Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("ZeroOneAI/ZEO-Med-2", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use ZeroOneAI/ZEO-Med-2 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "ZeroOneAI/ZEO-Med-2" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ZeroOneAI/ZEO-Med-2", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/ZeroOneAI/ZEO-Med-2
- SGLang
How to use ZeroOneAI/ZEO-Med-2 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "ZeroOneAI/ZEO-Med-2" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ZeroOneAI/ZEO-Med-2", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "ZeroOneAI/ZEO-Med-2" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ZeroOneAI/ZEO-Med-2", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use ZeroOneAI/ZEO-Med-2 with Docker Model Runner:
docker model run hf.co/ZeroOneAI/ZEO-Med-2
Access ZEO Med 2
ZEO Med 2, including the LoRA adapter weights, is released under the Apache License 2.0. Commercial use, modification and redistribution are permitted under that license. Access is granted automatically once you share your contact information.
By requesting access you confirm that you have read the Apache License 2.0 and the medical-use limitations described in this model card.
Log in or Sign Up to review the conditions and access this model content.
1. Model Introduction
ZEO = Zero ยท Expert ยท One. ZEO is ZeroOne AI's model family that brings expert-level intelligence into the journey from zero to one. ZEO Med is the medical-specialized ZEO model, designed to deliver expert-level medical reasoning in secure enterprise and hospital environments.
ZEO Med 2 is a medical language model for Korean healthcare licensing examinations and English medical question answering. On the Korean doctor licensing exam it is state of the art among open-weight models measured under the same evaluation setup, and on the nurse, pharmacist and dentist exams and on MedQA-USMLE it stays in the top tier.
ZEO Med 2 was developed by ZeroOne AI together with Samsung Medical Center, under the Advanced GPU Utilization Support Program of the Ministry of Science and ICT, Republic of Korea. See Acknowledgements.
2. Model Summary
| Model | ZEO Med 2 |
| Base model | google/gemma-4-31B-it |
| Adaptation | LoRA adapter (medical post-training) |
| Domain | Medical question answering |
| Languages | Korean, English |
| Answering | Single chat-completion call; internal reasoning, final answer scored |
| Decoding | temperature 0.7, top-p 0.95, 5 samples, majority vote |
| Max Tokens | 8,192 |
| Weights | LoRA adapter included (open weights) |
| License | Apache License 2.0 |
3. Evaluation Results
Every score below was measured under the same conditions. The model is never shown worked examples with their answers beforehand (0-shot), and each sample uses one prompt and one call in which the model reasons and then gives its final answer in the same response โ there is no second call asking for the answer. Each question is sampled five times and the most frequent answer is taken as the answer for that question. Blank cells were not measured under these conditions and are not filled in from other setups.
3.1 Open-weight models measured under the same setup
Doctor, Nurse, Pharmacist and Dentist are the Korean national licensing
examinations (KorMedMCQA). Every model in this table was measured by us under
the contract
final-sc5-medical-private-cot-final-answer-0shot-onecall-max8192-v1.
| Model | Doctor | Nurse | Pharmacist | Dentist | 4-exam average | MedQA-USMLE |
|---|---|---|---|---|---|---|
| ZEO Med 2 | 97.70% (425/435) | 96.36% (846/878) | 94.69% (838/885) | 87.05% (706/811) | 93.95% | 94.82% (1207/1273) |
| Gemma-4-31B-it | 97.01% (422/435) | 96.24% (845/878) | 95.03% (841/885) | 87.42% (709/811) | 93.93% | 94.74% (1206/1273) |
| Qwen3.6-27B | 93.56% (407/435) | 95.22% (836/878) | 93.90% (831/885) | 83.23% (675/811) | 91.48% | 94.11% (1198/1273) |
| Qwen3.6-35B-A3B | 93.10% (405/435) | 93.28% (819/878) | 92.77% (821/885) | 82.00% (665/811) | 90.29% | 94.34% (1201/1273) |
| Mistral Medium 3.5 ยถ | 90.80% (395/435) | 93.51% (821/878) | 92.99% (823/885) | 80.39% (652/811) | 89.42% | 91.99% (1171/1273) |
| HARI-Q2.5-Thinking * | 89.20% | 90.99% | 90.94% | 72.96% | 86.02% | 88.36% |
| gpt-oss-120b (native/default thinking) ยถ | 86.67% (377/435) | 89.07% (782/878) | 89.15% (789/885) | 75.83% (615/811) | 85.18% | 92.38% (1176/1273) |
| Nemotron 3 Nano | 83.22% (362/435) | 86.67% (761/878) | 85.76% (759/885) | 66.46% (539/811) | 80.53% | 88.45% (1126/1273) |
| Nemotron 3 Nano Omni | 81.84% (356/435) | 83.71% (735/878) | 88.02% (779/885) | 66.09% (536/811) | 79.92% | 83.03% (1057/1273) |
| MedGemma 1.0 | 72.41% (315/435) | 78.70% (691/878) | 75.37% (667/885) | 60.79% (493/811) | 71.82% | 86.65% (1103/1273) |
| MedGemma 1.5 | 67.59% (294/435) | 67.65% (594/878) | 67.23% (595/885) | 49.08% (398/811) | 62.89% | 72.43% (922/1273) |
Footnotes
- No mark: results we evaluated directly on open-weight/adapter models under the 0-shot SC@5 family of contracts. Per-model differences in runtime and parser are published in PROTOCOL.md.
*HARI: values published on its official model card. The four KorMedMCQA exams are 5-shot and MedQA-USMLE is 0-shot, andKorMed4 macrois the simple average of the published Doctor/Nurse/Pharmacist/Dentist scores.ยถMistral ยท gpt-oss: reference values from our own evaluation, run with a different reasoning setting from the rest of the table. Mistral Medium 3.5 ran with explicit high reasoning. gpt-oss-120b ran with its native/default thinking setting, which is why it is not labelledgpt-oss-120b (high).
3.2 KMed.ai
KMed.ai reported an average of 96.4 points on the 2025 Korean national doctor
licensing examination, in the
official announcement by Seoul National University Hospital and NAVER.
That is a score on the actual national examination, not on the KorMedMCQA
benchmark used throughout this model card. The two are different examinations
with different items, so the figure is not converted to %, no gap against
ZEO Med 2 is derived from it, and it is not placed in the table above.
3.3 Frontier API models
GPT-5.2 and Gemini 3.1 Flash-Lite are closed models reachable only through an
API, so they are listed separately from the open-weight models in
the Section 3.1 table. We
measured them ourselves through OpenRouter under the same evaluation conditions
as the Section 3.1 table.
Cells marked โ were not measured.
| Model | Doctor | Nurse | Pharmacist | Dentist | 4-exam average | MedQA-USMLE |
|---|---|---|---|---|---|---|
| ZEO Med 2 | 97.70% (425/435) | 96.36% (846/878) | 94.69% (838/885) | 87.05% (706/811) | 93.95% | 94.82% (1207/1273) |
| GPT-5.2 | 97.47% (424/435) | โ | โ | โ | โ | 95.99% (1222/1273) |
| Gemini 3.1 Flash-Lite | 96.09% (418/435) | 96.47% | 96.05% | 90.26% | 94.72% | 94.42% (1202/1273) |
The ZEO Med 2 row repeats the scores from the Section 3.1 table. ZEO Med 2 leads on the Korean doctor licensing examination. GPT-5.2 leads on MedQA-USMLE, and Gemini 3.1 Flash-Lite leads on the nurse, pharmacist and dentist examinations and on the four-exam average. This is a controlled comparison under matching evaluation conditions.
Gemini 3.1 Flash-Lite's nurse, pharmacist and dentist scores are recorded as percentages only, so no item counts are shown for those three cells.
4. Evaluation Protocol
Each question is answered in one chat-completion call. The model performs its reasoning inside that single call and then emits a final answer line; the scorer reads only that line. There is no separate rationale request and no second call, and hidden reasoning is neither scored nor released.
Prompting is template-driven. A system template fixes the answering format and a
user template carries the question, the option list and the option labels through
{question}, {choices} and {labels} placeholders. The exact templates used
for these scores are published here:
KorMedMCQA system ยท
KorMedMCQA user ยท
MedQA system ยท
MedQA user.
They hold placeholders only and carry no benchmark question text. Filled-in
prompts are not published.
Five samples are taken per question and the answer is the one most of them agree on. If a sample runs out of tokens before it finishes, that is recorded and published together with the score instead of being hidden.
The full setup is written out in PROTOCOL.md.
5. Run the Model
ZEO Med 2 is a LoRA adapter over google/gemma-4-31B-it. Serving loads the base
model and applies the adapter โ you do not merge weights. The adapter files are
in this repository under
evaluation-code/artifacts/adapter/ZEO-Med-31B-Adapter-v1/
(adapter_model.safetensors, adapter_config.json); access is gated with
automatic approval (see Section 10, License).
Serve with vLLM
The served model id is ZEO-Med-31B-Adapter-v1 (the identifier used throughout
the evaluation scripts); on Hugging Face the model is shown as ZEO Med 2.
python -m vllm.entrypoints.openai.api_server \
--model google/gemma-4-31B-it \
--enable-lora \
--lora-modules ZEO-Med-31B-Adapter-v1=evaluation-code/artifacts/adapter/ZEO-Med-31B-Adapter-v1 \
--max-lora-rank 8 \
--dtype bfloat16 \
--language-model-only \
--reasoning-parser gemma4 \
--max-model-len 16384
serve_vllm.sh (Section 6, Reproduce the Evaluation)
wraps this exact command with the pinned base
revision and the adapter hash check.
Chat
from openai import OpenAI
client = OpenAI(base_url="http://localhost:8000/v1", api_key="EMPTY")
resp = client.chat.completions.create(
model="ZEO-Med-31B-Adapter-v1",
messages=[
{"role": "user", "content": "๊ณ ํ์ 1์ฐจ ์ฝ์ ์ ํ์ ์ผ๋ฐ์ ์์น์ ์ค๋ช
ํด์ค."},
],
temperature=0.7,
top_p=0.95,
max_tokens=8192,
)
print(resp.choices[0].message.content)
The model reasons internally within the single call and returns the final answer
in content. The base model's tokenizer and chat template are used; the adapter
does not replace them.
The model is for research and non-clinical use. See Sections 7 and 8.
6. Reproduce the Evaluation
This is separate from Section 5, Run the Model: here you re-run the exact 0-shot SC@5 contract to reproduce the published scores. Benchmark question text is not redistributed โ bring your own licensed copy of the datasets. Full step-by-step instructions, including the CPU-only package self-test, are in the evaluation-code guide.
run_full5_reproduction.sh does not start vLLM itself: you first serve the five
task endpoints (with serve_vllm.sh zeo), then pass their URLs, PIDs, and the
input/manifest directories as environment variables. It requires all of the
following to be set:
export PACKAGE_ROOT=/path/to/evaluation-code
export MODEL_PACKAGE_ROOT=/path/to/evaluation-code
export SERVE_PY=/path/to/vllm-0.23.0/bin/python
export SERVED_MODEL=ZEO-Med-31B-Adapter-v1
export SERVING_MODE=dynamic_lora
export FULL5_INPUT_DIR=/path/to/full5-inputs # your licensed benchmark rows
export FULL5_MANIFEST_DIR=/path/to/full5-manifests
export OUTPUT_ROOT=/path/to/new-full5-output
export NUM_WORKERS=48
export REQUEST_SEED=42
# one URL + PID per task, from the endpoints you started with serve_vllm.sh:
export DOCTOR_URL=... NURSE_URL=... PHARMACIST_URL=... DENTIST_URL=... MEDQA_URL=...
export DOCTOR_SERVE_PID=... NURSE_SERVE_PID=... PHARMACIST_SERVE_PID=... DENTIST_SERVE_PID=... MEDQA_SERVE_PID=...
bash evaluation-code/scripts/run_full5_reproduction.sh
A full-5 macro delta of zero and an all-task match indicate a successful reproduction.
7. Evaluation Datasets
The scores above were measured on the public test sets below. Because of the licenses of those test sets, no question text or dataset file is redistributed here; anyone reproducing the measurement obtains them under their own terms.
- KorMedMCQA โ doctor,
nurse, pharmacist and dentist configurations,
testsplit, revision79efd6f91edfc8036330d7a4daa88b9f2deb9a82. The dataset card statesCC-BY-NC-2.0. - MedQA-USMLE-4-options
โ
testsplit, revision0fb93dd23a7339b6dcd27e241cb9b5eca62d4d18. The dataset card statesCC-BY-4.0.
Question text is excluded from this repository by publication policy. This is a release-control decision, not legal advice.
8. Intended Use and Limitations
ZEO Med 2 is released for research, education, evaluation, and product development. It is not a medical device and is not approved for autonomous diagnosis, prescription, treatment, or other independent clinical decision-making, and may not be used as the sole basis for a clinical decision. Any healthcare deployment must include qualified professional oversight, institution-specific validation, and compliance with applicable laws.
The published scores measure exam-style multiple-choice answer selection under one fixed setup and do not establish clinical safety or diagnostic accuracy.
9. Base Model and Modifications
ZEO Med 2 is a modified and fine-tuned derivative of:
- Base model:
google/gemma-4-31B-it(revision3548789868c5356dbf307c98e6f609007b82b3eb) - Base model license: Apache License 2.0 (Developer: Google)
- Modifications by ZeroOne AI: medical-domain LoRA fine-tuning, Korean/English medical instruction tuning, response alignment, and configuration changes
ZEO Med 2 is not affiliated with, endorsed by, or sponsored by Google. The full
license text is in LICENSE; attribution and modification notices are
in NOTICE.
10. License
ZEO Med 2 is released under the Apache License 2.0, and that includes the LoRA adapter weights. Commercial use, modification, and redistribution are permitted under that license.
The base model google/gemma-4-31B-it is under the same license. Attribution and
the list of modifications made by ZeroOne AI are in NOTICE, which you
must retain when redistributing the model or a derivative of it, as required by
Section 4 of the license.
Access to the repository files is gated with automatic approval: you share your contact information once and access is granted immediately.
11. Contact
- Email: zeo@zeroone.ai
- Company: zeroone.ai
- Product: AInode
12. Acknowledgements
ZEO Med 2 was developed under the Advanced GPU Utilization Support Program of the
Ministry of Science and ICT, Republic of Korea, project no. 02-26-01-0282,
awarded to Samsung Medical Center.
We thank the research team of the Department of Emergency Medicine, Samsung Medical Center โ Prof. Won Chul Cha, principal investigator, and Prof. Meong Hi Son โ for the collaboration on this project.
13. Citation
@misc{zeo-med-2-2026,
title={ZEO Med 2},
author={ZeroOne AI},
year={2026},
url={https://huggingface.co/ZeroOneAI/ZEO-Med-2}
}
ํ๊ตญ์ด
1. ๋ชจ๋ธ ์๊ฐ
ZEO = Zero ยท Expert ยท One. ZEO๋ (์ฃผ)์ ๋ก์์์ด์์ด(ZeroOne AI)์ ์์ฒด AI ๋ชจ๋ธ ํจ๋ฐ๋ฆฌ๋ก, Zero์์ One์ผ๋ก ์ด์ด์ง๋ ๋ฌธ์ ํด๊ฒฐ ๊ณผ์ ์ ๋๋ฉ์ธ ์ ๋ฌธ๊ฐ ์์ค์ ์ง๋ฅ์ ๊ฒฐํฉํ๋ค๋ ์๋ฏธ๋ฅผ ๋ด๊ณ ์์ต๋๋ค. ZEO Med๋ ์๋ฃ ๋ถ์ผ์ ํนํ๋ ZEO ๋ชจ๋ธ๋ก, ์์ ํ ๊ธฐ์ ๋ฐ ๋ณ์ ํ๊ฒฝ์์ ์ ๋ฌธ๊ฐ ์์ค์ ์ํ์ ์ถ๋ก ์ ์ ๊ณตํ๋๋ก ์ค๊ณ๋์์ต๋๋ค.
์ง์ค๋ฉ๋2(ZEO Med 2)๋ ํ๊ตญ ์๋ฃ ๊ตญ๊ฐ๊ณ ์์ ์๋ฌธ ์๋ฃ ๋ฌธ์ ํ์ด๋ฅผ ์ํ ์๋ฃ ์ธ์ด ๋ชจ๋ธ์ ๋๋ค. ์์ฌ ๊ตญ๊ฐ๊ณ ์์์๋ ๋์ผํ ํ๊ฐ ํ๊ฒฝ์ผ๋ก ์ธก์ ํ open-weight ๋ชจ๋ธ ์ค ์ต๊ณ ์ฑ๋ฅ์ด๋ฉฐ, ๊ฐํธ์ฌยท์ฝ์ฌยท์น๊ณผ์์ฌ ๊ตญ๊ฐ๊ณ ์์ MedQA-USMLE์์๋ ์ต์์๊ถ ์ฑ์ ์ ์ ์งํฉ๋๋ค.
ZEO Med 2๋ ๊ณผํ๊ธฐ์ ์ ๋ณดํต์ ๋ถ ์ฒจ๋จ GPU ํ์ฉ ์ง์ ์ฌ์ ์ ์ง์์ ๋ฐ์, (์ฃผ)์ ๋ก์์์ด์์ด(ZeroOne AI)๊ฐ ์ผ์ฑ์์ธ๋ณ์๊ณผ ํจ๊ป ๊ฐ๋ฐํ์ต๋๋ค. ์์ธํ ๋ด์ฉ์ ๊ฐ์ฌ์ ๊ธ์ ์์ต๋๋ค.
2. ๋ชจ๋ธ ๊ฐ์
| ๋ชจ๋ธ | ์ง์ค๋ฉ๋2(ZEO Med 2) |
| ๊ธฐ๋ฐ ๋ชจ๋ธ | google/gemma-4-31B-it |
| ์ ์ ๋ฐฉ์ | LoRA ์ด๋ํฐ (์๋ฃ ์ฌํํ์ต) |
| ๋ถ์ผ | ์๋ฃ ๋ฌธ์ ํ์ด |
| ์ธ์ด | ํ๊ตญ์ด, ์์ด |
| ์๋ต ๋ฐฉ์ | ํ ๋ฒ์ ํธ์ถ๋ก ์ถ๋ก ๊ณผ ์ต์ข ๋ต์ ํจ๊ป ์์ฑ, ์ต์ข ๋ต๋ง ์ฑ์ |
| ๋์ฝ๋ฉ | temperature 0.7, top-p 0.95, 5ํ ์ํ๋ง ํ ๋ค์๊ฒฐ |
| ์ต๋ ํ ํฐ | 8,192 |
| ๊ฐ์ค์น | LoRA ์ด๋ํฐ ํฌํจ (์คํ ์จ์ดํธ) |
| ๋ผ์ด์ผ์ค | Apache License 2.0 |
3. ํ๊ฐ ๊ฒฐ๊ณผ
์๋ ์ ์๋ ๋ชจ๋ ๊ฐ์ ์กฐ๊ฑด์์ ์ธก์ ํ์ต๋๋ค. ๋ชจ๋ธ์ ์์ ๋ฌธํญ๊ณผ ์ ๋ต์ ๋ฏธ๋ฆฌ ๋ณด์ฌ์ฃผ์ง ์์ผ๋ฉฐ(0-shot), ํ๋์ ํ๋กฌํํธ๋ก ํ ๋ฒ ํธ์ถํด ๋ชจ๋ธ์ด ์ถ๋ก ์ ๊ฑฐ์น ๋ค ์ต์ข ๋ต์ ๊ฐ์ ์๋ต ์์์ ์ด์ด์ ๋ด๋๊ฒ ํฉ๋๋ค. ๋ต์ ๋ค์ ๋ฌป๋ ๋ ๋ฒ์งธ ํธ์ถ์ ์์ต๋๋ค. ๋ฌธํญ๋ง๋ค 5ํ ์ํ๋งํด ๊ฐ์ฅ ๋ง์ด ๋์จ ๋ต์ ๊ทธ ๋ฌธํญ์ ๋ต์ผ๋ก ์ฑํํฉ๋๋ค. ๋น์นธ์ ์ด ์กฐ๊ฑด์ผ๋ก ์ธก์ ํ์ง ์์ ๊ฐ์ด๋ฉฐ, ๋ค๋ฅธ ์กฐ๊ฑด์ ๊ฐ์ผ๋ก ์ฑ์ฐ์ง ์์์ต๋๋ค.
3.1 ๊ฐ์ ์กฐ๊ฑด์ผ๋ก ์ธก์ ํ open-weight ๋ชจ๋ธ
DoctorยทNurseยทPharmacistยทDentist๋ ํ๊ตญ ์๋ฃ ๊ตญ๊ฐ๊ณ ์(KorMedMCQA)์
๋๋ค. ์ด ํ์
๋ชจ๋ ๋ชจ๋ธ์ ๋น์ฌ๊ฐ
final-sc5-medical-private-cot-final-answer-0shot-onecall-max8192-v1 ๊ณ์ฝ์ผ๋ก
์ง์ ์ธก์ ํ์ต๋๋ค.
| ๋ชจ๋ธ | Doctor | Nurse | Pharmacist | Dentist | 4๊ฐ ์ง๊ตฐ ํ๊ท | MedQA-USMLE |
|---|---|---|---|---|---|---|
| ZEO Med 2 | 97.70% (425/435) | 96.36% (846/878) | 94.69% (838/885) | 87.05% (706/811) | 93.95% | 94.82% (1207/1273) |
| Gemma-4-31B-it | 97.01% (422/435) | 96.24% (845/878) | 95.03% (841/885) | 87.42% (709/811) | 93.93% | 94.74% (1206/1273) |
| Qwen3.6-27B | 93.56% (407/435) | 95.22% (836/878) | 93.90% (831/885) | 83.23% (675/811) | 91.48% | 94.11% (1198/1273) |
| Qwen3.6-35B-A3B | 93.10% (405/435) | 93.28% (819/878) | 92.77% (821/885) | 82.00% (665/811) | 90.29% | 94.34% (1201/1273) |
| Mistral Medium 3.5 ยถ | 90.80% (395/435) | 93.51% (821/878) | 92.99% (823/885) | 80.39% (652/811) | 89.42% | 91.99% (1171/1273) |
| HARI-Q2.5-Thinking * | 89.20% | 90.99% | 90.94% | 72.96% | 86.02% | 88.36% |
| gpt-oss-120b (native/default thinking) ยถ | 86.67% (377/435) | 89.07% (782/878) | 89.15% (789/885) | 75.83% (615/811) | 85.18% | 92.38% (1176/1273) |
| Nemotron 3 Nano | 83.22% (362/435) | 86.67% (761/878) | 85.76% (759/885) | 66.46% (539/811) | 80.53% | 88.45% (1126/1273) |
| Nemotron 3 Nano Omni | 81.84% (356/435) | 83.71% (735/878) | 88.02% (779/885) | 66.09% (536/811) | 79.92% | 83.03% (1057/1273) |
| MedGemma 1.0 | 72.41% (315/435) | 78.70% (691/878) | 75.37% (667/885) | 60.79% (493/811) | 71.82% | 86.65% (1103/1273) |
| MedGemma 1.5 | 67.59% (294/435) | 67.65% (594/878) | 67.23% (595/885) | 49.08% (398/811) | 62.89% | 72.43% (922/1273) |
ํ์๊ณผ ๊ฐ์ฃผ
- ํ์ ์์: ๋น์ฌ๊ฐ 0-shot SC@5 ๊ณ์ด ๊ณ์ฝ์ผ๋ก ์ง์ ํ๊ฐํ open-weight/adapter ๊ฒฐ๊ณผ์ ๋๋ค. ๋ชจ๋ธ๋ณ runtimeยทparser ์ฐจ์ด๋ PROTOCOL.md์์ ๊ณต๊ฐํฉ๋๋ค.
*HARI: ๊ณต์ ๋ชจ๋ธ์นด๋ ๊ณต๊ฐ๊ฐ์ ๋๋ค. KorMedMCQA 4๊ฐ ์ํ์ 5-shot, MedQA-USMLE๋ 0-shot์ด๋ฉฐ,KorMed4 macro๋ ๊ณต๊ฐ๋ Doctor/Nurse/Pharmacist/Dentist ์ ์์ ๋จ์ ํ๊ท ์ ๋๋ค.ยถMistralยทgpt-oss: ๋น์ฌ ์ง์ ํ๊ฐ ์ฐธ๊ณ ๊ฐ์ด๋ฉฐ, ํ์ ๋๋จธ์ง ๋ชจ๋ธ๊ณผ ์ถ๋ก ์ค์ ์ด ๋ค๋ฆ ๋๋ค. Mistral Medium 3.5๋ explicit high reasoning์ผ๋ก, gpt-oss-120b๋ native/default thinking ์ค์ ์ผ๋ก ์ธก์ ํ์ต๋๋ค. ๊ทธ๋์ ํ์ฌ ๊ฐ์gpt-oss-120b (high)๋ก ํ๊ธฐํ์ง ์์ต๋๋ค.
3.2 KMed.ai
KMed.ai๋
์์ธ๋๋ณ์ยท๋ค์ด๋ฒ ๊ณต์ ๋ฐํ์์
2025๋
๋ ์์ฌ ๊ตญ๊ฐ๊ณ ์ ํ๊ท 96.4์ ์ ๊ณต๊ฐํ์ต๋๋ค. ์ด๋ ์ค์ ๊ตญ๊ฐ๊ณ ์์์ ๋ฐ์
์ ์์ด๋ฉฐ, ์ด ๋ชจ๋ธ์นด๋ ์ ๋ฐ์์ ์ฐ๋ KorMedMCQA ๋ฒค์น๋งํฌ ์ ์๊ฐ ์๋๋๋ค. ๋
์ํ์ ๋ฌธํญ์ด ๋ค๋ฅธ ๋ณ๊ฐ์ ์ํ์ด๋ฏ๋ก ์ด ๊ฐ์ %๋ก ๋ฐ๊พธ์ง ์๊ณ , ZEO Med 2์์
์ ์ ์ฐจ์ด๋ ๊ณ์ฐํ์ง ์์ผ๋ฉฐ, ์ ํ์๋ ๋ฃ์ง ์์ต๋๋ค.
3.3 ํ๋ก ํฐ์ด API ๋ชจ๋ธ
GPT-5.2์ Gemini 3.1 Flash-Lite๋ API๋ก๋ง ์ ๊ทผํ ์ ์๋ ๋น๊ณต๊ฐ ๋ชจ๋ธ์ด๋ผ,
3.1์ ํ์ open-weight ๋ชจ๋ธ๊ณผ ๋ถ๋ฆฌํด
์ฃ์ต๋๋ค. ๋น์ฌ๊ฐ OpenRouter๋ฅผ ํตํด
3.1์ ํ์ ๊ฐ์ ํ๊ฐ ์กฐ๊ฑด์ผ๋ก ์ง์
์ธก์ ํ์ต๋๋ค. โ๋ ์ธก์ ํ์ง ์์ ๊ฐ์
๋๋ค.
| ๋ชจ๋ธ | Doctor | Nurse | Pharmacist | Dentist | 4๊ฐ ์ง๊ตฐ ํ๊ท | MedQA-USMLE |
|---|---|---|---|---|---|---|
| ZEO Med 2 | 97.70% (425/435) | 96.36% (846/878) | 94.69% (838/885) | 87.05% (706/811) | 93.95% | 94.82% (1207/1273) |
| GPT-5.2 | 97.47% (424/435) | โ | โ | โ | โ | 95.99% (1222/1273) |
| Gemini 3.1 Flash-Lite | 96.09% (418/435) | 96.47% | 96.05% | 90.26% | 94.72% | 94.42% (1202/1273) |
ZEO Med 2 ํ์ 3.1์ ํ์ ์ ์๋ฅผ ๋ค์ ์ฎ๊ธด ๊ฒ์ ๋๋ค. ์์ฌ ๊ตญ๊ฐ๊ณ ์๋ ZEO Med 2๊ฐ, MedQA-USMLE๋ GPT-5.2๊ฐ, ๊ฐํธ์ฌยท์ฝ์ฌยท์น๊ณผ์์ฌ ๊ตญ๊ฐ๊ณ ์์ 4๊ฐ ์ง๊ตฐ ํ๊ท ์ Gemini 3.1 Flash-Lite๊ฐ ๊ฐ์ฅ ๋์ต๋๋ค. ํ๊ฐ ์กฐ๊ฑด์ด ์๋ก ๊ฐ์ ํต์ ๋ ๋น๊ต ์งํ์ ๋๋ค.
Gemini 3.1 Flash-Lite์ ๊ฐํธ์ฌยท์ฝ์ฌยท์น๊ณผ์์ฌ ์ ์๋ ๋ฐฑ๋ถ์จ๋ก๋ง ๊ธฐ๋ก๋์ด ์์ด ๊ทธ ์ธ ์นธ์๋ ๋ฌธํญ ์๋ฅผ ํ๊ธฐํ์ง ์์์ต๋๋ค.
4. ํ๊ฐ ๋ฐฉ๋ฒ
๋ฌธํญ๋ง๋ค ํ ๋ฒ์ chat-completion ํธ์ถ๋ก ๋ตํฉ๋๋ค. ๋ชจ๋ธ์ ๊ทธ ํธ์ถ ์์์ ์ถ๋ก ์ ์ํํ ๋ค ์ต์ข ๋ต ํ ์ค์ ๋ด๋๊ณ , ์ฑ์ ๊ธฐ๋ ๊ทธ ์ค๋ง ์ฝ์ต๋๋ค. ๊ทผ๊ฑฐ๋ฅผ ๋ฐ๋ก ์์ฒญํ๋ ์ ์ฐจ๋ ๋ ๋ฒ์งธ ํธ์ถ์ ์์ผ๋ฉฐ, ๋ด๋ถ ์ถ๋ก ์ ์ฑ์ ํ์ง๋ ๊ณต๊ฐํ์ง๋ ์์ต๋๋ค.
ํ๋กฌํํธ๋ ํ
ํ๋ฆฟ์ผ๋ก ๊ตฌ์ฑํฉ๋๋ค. system ํ
ํ๋ฆฟ์ด ๋ต๋ณ ํ์์ ๊ณ ์ ํ๊ณ , user
ํ
ํ๋ฆฟ์ด {question}, {choices}, {labels} ์๋ฆฌํ์์๋ก ๋ฌธํญ๊ณผ ์ ํ์ง,
์ ํ์ง ๊ธฐํธ๋ฅผ ์ ๋ฌํฉ๋๋ค. ์ด ์ ์๋ฅผ ๋ผ ๋ ์ฌ์ฉํ ์ค์ ํ
ํ๋ฆฟ์ ๊ณต๊ฐํฉ๋๋ค:
KorMedMCQA system ยท
KorMedMCQA user ยท
MedQA system ยท
MedQA user.
์๋ฆฌํ์์๋ง ๋ค์ด ์๊ณ ๋ฒค์น๋งํฌ ๋ฌธํญ ์๋ฌธ์ ํฌํจํ์ง ์์ต๋๋ค. ๊ฐ์ด ์ฑ์์ง
ํ๋กฌํํธ๋ ๊ณต๊ฐํ์ง ์์ต๋๋ค.
๋ฌธํญ๋ง๋ค 5ํ๋ฅผ ์ํ๋งํด ๊ฐ์ฅ ๋ง์ด ๋์จ ๋ต์ ์ฑํํฉ๋๋ค. ์ด๋ค ์ํ์ด ํ ํฐ์ ๋ค ์ฐ๊ณ ๋๋๋ฉด ๊ทธ ์ฌ์ค๋ ์ ์์ ํจ๊ป ๊ณต๊ฐํฉ๋๋ค.
์ ์ฒด ์ค์ ์ PROTOCOL.md์ ์ ๋ฆฌ๋ผ ์์ต๋๋ค.
5. ๋ชจ๋ธ ์ฌ์ฉ
ZEO Med 2๋ google/gemma-4-31B-it ์์ ์น๋ LoRA ์ด๋ํฐ์
๋๋ค. ์๋น ์ ๊ธฐ๋ฐ
๋ชจ๋ธ์ ๋ก๋ํ๊ณ ์ด๋ํฐ๋ฅผ ์ ์ฉํ๋ฉฐ, ๊ฐ์ค์น๋ฅผ ๋ณํฉํ์ง ์์ต๋๋ค. ์ด๋ํฐ ํ์ผ์ ์ด
์ ์ฅ์์ evaluation-code/artifacts/adapter/ZEO-Med-31B-Adapter-v1/
(adapter_model.safetensors, adapter_config.json)์ ์๊ณ , ์ ๊ทผ์ ์๋ ์น์ธ
๋ฐฉ์์ gated์
๋๋ค(10์ ๋ผ์ด์ผ์ค ์ฐธ๊ณ ).
vLLM ์๋น
์๋น ๋ชจ๋ธ id๋ ZEO-Med-31B-Adapter-v1์
๋๋ค(ํ๊ฐ ์คํฌ๋ฆฝํธ ์ ๋ฐ์์ ์ฐ๋
์๋ณ์). Hugging Face ํ์ด์ง์๋ ZEO Med 2๋ก ํ์๋ฉ๋๋ค.
python -m vllm.entrypoints.openai.api_server \
--model google/gemma-4-31B-it \
--enable-lora \
--lora-modules ZEO-Med-31B-Adapter-v1=evaluation-code/artifacts/adapter/ZEO-Med-31B-Adapter-v1 \
--max-lora-rank 8 \
--dtype bfloat16 \
--language-model-only \
--reasoning-parser gemma4 \
--max-model-len 16384
serve_vllm.sh(6์ ํ๊ฐ ์ฌํ)๊ฐ ๊ธฐ๋ฐ revision ๊ณ ์ ๊ณผ ์ด๋ํฐ ํด์ ๊ฒ์ฆ๊น์ง ํฌํจํด ์
๋ช
๋ น์ ๊ทธ๋๋ก ๊ฐ์๋๋ค.
๋ํ
from openai import OpenAI
client = OpenAI(base_url="http://localhost:8000/v1", api_key="EMPTY")
resp = client.chat.completions.create(
model="ZEO-Med-31B-Adapter-v1",
messages=[
{"role": "user", "content": "๊ณ ํ์ 1์ฐจ ์ฝ์ ์ ํ์ ์ผ๋ฐ์ ์์น์ ์ค๋ช
ํด์ค."},
],
temperature=0.7,
top_p=0.95,
max_tokens=8192,
)
print(resp.choices[0].message.content)
๋ชจ๋ธ์ ํ ๋ฒ์ ํธ์ถ ์์์ ์ถ๋ก ์ ์ํํ๊ณ ์ต์ข
๋ต์ content๋ก ๋ฐํํฉ๋๋ค.
ํ ํฌ๋์ด์ ์ chat template์ ๊ธฐ๋ฐ ๋ชจ๋ธ์ ๊ฒ์ ์ฌ์ฉํ๋ฉฐ, ์ด๋ํฐ๊ฐ ์ด๋ฅผ ๋ฐ๊พธ์ง
์์ต๋๋ค.
์ด ๋ชจ๋ธ์ ์ฐ๊ตฌยท๋น์์ ์ฉ๋์ ๋๋ค. 7์ ํ๊ฐ ๋ฐ์ดํฐ์ ๊ณผ 8์ ์ฌ์ฉ ๋ชฉ์ ๊ณผ ํ๊ณ๋ฅผ ์ฐธ๊ณ ํ์ญ์์ค.
6. ํ๊ฐ ์ฌํ
5์ ๋ชจ๋ธ ์ฌ์ฉ๊ณผ ๋ณ๊ฐ์ ๋๋ค. ์ฌ๊ธฐ์๋ ๊ณต๊ฐ๋ ์ ์๋ฅผ ์ฌํํ๊ธฐ ์ํด 0-shot SC@5 ๊ณ์ฝ์ ๊ทธ๋๋ก ๋ค์ ์คํํฉ๋๋ค. ๋ฒค์น๋งํฌ ๋ฌธํญ ์๋ฌธ์ ์ฌ๋ฐฐํฌํ์ง ์์ผ๋ ๊ฐ์ ๋ผ์ด์ผ์ค์ ๋ฐ๋ผ ๋ฐ์ดํฐ์ ์ ์ง์ ๋ฐ์์ผ ํฉ๋๋ค. CPU ์ ์ฉ ํจํค์ง self-test๋ฅผ ํฌํจํ ๋จ๊ณ๋ณ ์์ธ ์ ์ฐจ๋ ํ๊ฐ ์ฝ๋ ์๋ด์ ์์ต๋๋ค.
run_full5_reproduction.sh๋ vLLM์ ์ง์ ๋์ฐ์ง ์์ต๋๋ค. ๋จผ์ 5๊ฐ ์ํ
์๋ํฌ์ธํธ๋ฅผ ์๋น(serve_vllm.sh zeo)ํ ๋ค, ๊ทธ URLยทPID์ ์
๋ ฅ/๋งค๋ํ์คํธ
๋๋ ํฐ๋ฆฌ๋ฅผ ํ๊ฒฝ๋ณ์๋ก ๋๊ฒจ์ผ ํฉ๋๋ค. ๋ค์์ ๋ชจ๋ ์ค์ ํด์ผ ํฉ๋๋ค.
export PACKAGE_ROOT=/path/to/evaluation-code
export MODEL_PACKAGE_ROOT=/path/to/evaluation-code
export SERVE_PY=/path/to/vllm-0.23.0/bin/python
export SERVED_MODEL=ZEO-Med-31B-Adapter-v1
export SERVING_MODE=dynamic_lora
export FULL5_INPUT_DIR=/path/to/full5-inputs # ๊ฐ์ ๋ผ์ด์ผ์ค๋ก ๋ฐ์ ๋ฒค์น๋งํฌ ๋ฌธํญ
export FULL5_MANIFEST_DIR=/path/to/full5-manifests
export OUTPUT_ROOT=/path/to/new-full5-output
export NUM_WORKERS=48
export REQUEST_SEED=42
# serve_vllm.sh๋ก ๋์ด ์๋ํฌ์ธํธ์ ์ํ๋ณ URL + PID:
export DOCTOR_URL=... NURSE_URL=... PHARMACIST_URL=... DENTIST_URL=... MEDQA_URL=...
export DOCTOR_SERVE_PID=... NURSE_SERVE_PID=... PHARMACIST_SERVE_PID=... DENTIST_SERVE_PID=... MEDQA_SERVE_PID=...
bash evaluation-code/scripts/run_full5_reproduction.sh
5๊ฐ ์ํ macro delta๊ฐ 0์ด๊ณ ๋ชจ๋ ์ํ์ด ์ผ์นํ๋ฉด ์ฌํ์ ์ฑ๊ณตํ ๊ฒ์ ๋๋ค.
7. ํ๊ฐ ๋ฐ์ดํฐ์
์ ์ ์๋ ์๋ ๊ณต๊ฐ ํ ์คํธ์ ์ผ๋ก ์ธก์ ํ์ต๋๋ค. ๊ฐ ํ ์คํธ์ ์ ๋ผ์ด์ผ์ค ๋๋ฌธ์ ๋ฌธํญ ์๋ฌธ์ด๋ ๋ฐ์ดํฐ์ ํ์ผ์ ์ด ์ ์ฅ์์ ์ฌ๋ฐฐํฌํ์ง ์์ผ๋ฉฐ, ์ฌํํ๋ ค๋ ์ชฝ์์ ๊ฐ์์ ์ด์ฉ ์กฐ๊ฑด์ ๋ฐ๋ผ ์ง์ ๋ฐ์์ผ ํฉ๋๋ค.
- KorMedMCQA โ doctor,
nurse, pharmacist, dentist ๊ตฌ์ฑ,
testsplit, revision79efd6f91edfc8036330d7a4daa88b9f2deb9a82. ๋ฐ์ดํฐ์ ์นด๋ ํ๊ธฐ๋CC-BY-NC-2.0์ ๋๋ค. - MedQA-USMLE-4-options
โ
testsplit, revision0fb93dd23a7339b6dcd27e241cb9b5eca62d4d18. ๋ฐ์ดํฐ์ ์นด๋ ํ๊ธฐ๋CC-BY-4.0์ ๋๋ค.
๋ฌธํญ ์๋ฌธ์ ๊ณต๊ฐ ์ ์ฑ ์ ์ด ์ ์ฅ์์์ ์ ์ธํ์ต๋๋ค. ์ด๋ ๋ฐฐํฌ ํต์ ๊ฒฐ์ ์ด๋ฉฐ ๋ฒ๋ฅ ์๋ฌธ์ด ์๋๋๋ค.
8. ์ฌ์ฉ ๋ชฉ์ ๊ณผ ํ๊ณ
ZEO Med 2๋ ์ฐ๊ตฌยท๊ต์กยทํ๊ฐยท์ ํ ๊ฐ๋ฐ์ ์ํด ๊ณต๊ฐํ ๋ชจ๋ธ์ ๋๋ค. ์๋ฃ๊ธฐ๊ธฐ๊ฐ ์๋๋ฉฐ ์์จ์ ์ง๋จยท์ฒ๋ฐฉยท์น๋ฃ ๋ฑ ๋ ๋ฆฝ์ ์์ ์์ฌ๊ฒฐ์ ์ฉ์ผ๋ก ์น์ธ๋์ง ์์๊ณ , ์์ ํ๋จ์ ์ ์ผํ ๊ทผ๊ฑฐ๋ก ์ฌ์ฉํ ์ ์์ต๋๋ค. ์๋ฃ ํ์ฅ ๋์ ์์๋ ์๊ฒฉ ์๋ ์ ๋ฌธ๊ฐ์ ๊ฐ๋ , ๊ธฐ๊ด๋ณ ๊ฒ์ฆ, ๊ด๋ จ ๋ฒ๊ท ์ค์๊ฐ ํ์ํฉ๋๋ค.
๊ณต๊ฐ๋ ์ ์๋ ๊ณ ์ ๋ ์กฐ๊ฑด์์์ ์ํํ ๊ฐ๊ด์ ์ ๋ต ์ ํ์ ์ธก์ ํ ๊ฒ์ผ๋ก, ์์์ ์์ ์ฑ์ด๋ ์ง๋จ ์ ํ๋๋ฅผ ์ ์ฆํ์ง ์์ต๋๋ค.
9. ๊ธฐ๋ฐ ๋ชจ๋ธ๊ณผ ์์ ์ฌํญ
ZEO Med 2๋ ๋ค์์ ์์ ยทํ์ธํ๋ํ ํ์ ๋ชจ๋ธ์ ๋๋ค.
- ๊ธฐ๋ฐ ๋ชจ๋ธ:
google/gemma-4-31B-it(revision3548789868c5356dbf307c98e6f609007b82b3eb) - ๊ธฐ๋ฐ ๋ชจ๋ธ ๋ผ์ด์ผ์ค: Apache License 2.0 (๊ฐ๋ฐ: Google)
- (์ฃผ)์ ๋ก์์์ด์์ด(ZeroOne AI)์ ์์ : ์๋ฃ ๋๋ฉ์ธ LoRA ํ์ธํ๋, ํ๊ตญ์ดยท์์ด ์๋ฃ instruction ํ๋, ์๋ต ์ ๋ ฌ, ๊ตฌ์ฑ ๋ณ๊ฒฝ
ZEO Med 2๋ Google๊ณผ ์ ํดยท๋ณด์ฆยทํ์ ๊ด๊ณ๊ฐ ์์ต๋๋ค. ๋ผ์ด์ผ์ค ์ ๋ฌธ์
LICENSE์, ๊ท์ยท์์ ๊ณ ์ง๋ NOTICE์ ์์ต๋๋ค.
10. ๋ผ์ด์ผ์ค
ZEO Med 2๋ LoRA ์ด๋ํฐ ๊ฐ์ค์น๋ฅผ ํฌํจํด Apache License 2.0์ผ๋ก ๊ณต๊ฐํฉ๋๋ค. ์ด ๋ผ์ด์ผ์ค์ ๋ฐ๋ผ ์์ ์ ์ด์ฉ, ์์ , ์ฌ๋ฐฐํฌ๊ฐ ๋ชจ๋ ํ์ฉ๋ฉ๋๋ค.
๊ธฐ๋ฐ ๋ชจ๋ธ google/gemma-4-31B-it๋ ๊ฐ์ ๋ผ์ด์ผ์ค์
๋๋ค. ๊ท์ ํ๊ธฐ์
(์ฃผ)์ ๋ก์์์ด์์ด(ZeroOne AI)๊ฐ
๊ฐํ ์์ ๋ชฉ๋ก์ NOTICE์ ์์ผ๋ฉฐ, ๋ผ์ด์ผ์ค ์ 4์กฐ์ ๋ฐ๋ผ ์ด ๋ชจ๋ธ์ด๋
ํ์ ๋ชจ๋ธ์ ์ฌ๋ฐฐํฌํ ๋ ํจ๊ป ์ ์งํด์ผ ํฉ๋๋ค.
์ ์ฅ์ ํ์ผ ์ ๊ทผ์ ์๋ ์น์ธ ๋ฐฉ์์ gated์ ๋๋ค. ์ฐ๋ฝ์ฒ๋ฅผ ํ ๋ฒ ์ ์ถํ๋ฉด ์ฆ์ ์ ๊ทผ ๊ถํ์ด ๋ถ์ฌ๋ฉ๋๋ค.
11. ๋ฌธ์
- ์ด๋ฉ์ผ: zeo@zeroone.ai
- ํ์ฌ ์๊ฐ: zeroone.ai
- ์ ํ: AInode
12. ๊ฐ์ฌ์ ๊ธ
ZEO Med 2๋ ๊ณผํ๊ธฐ์ ์ ๋ณดํต์ ๋ถ ์ฒจ๋จ GPU ํ์ฉ ์ง์ ์ฌ์
(๊ณผ์ ๋ฒํธ
02-26-01-0282, ์ฃผ๊ด ์ผ์ฑ์์ธ๋ณ์)์ ์ง์์ ๋ฐ์ ๊ฐ๋ฐํ์ต๋๋ค.
ํจ๊ป ์ฐ๊ตฌ๋ฅผ ์งํํด ์ฃผ์ ์ผ์ฑ์์ธ๋ณ์ ์๊ธ์ํ๊ณผ ์ฐจ์์ฒ ๊ต์(์ฐ๊ตฌ์ฑ ์์)์ ์๋ช ํฌ ๊ต์ ์ฐ๊ตฌํ์ ๊ฐ์ฌ๋๋ฆฝ๋๋ค.
13. Citation
@misc{zeo-med-2-2026,
title={ZEO Med 2},
author={ZeroOne AI},
year={2026},
url={https://huggingface.co/ZeroOneAI/ZEO-Med-2}
}

