Instructions to use DataXAI/AnesTRACE-Eval with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use DataXAI/AnesTRACE-Eval with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="DataXAI/AnesTRACE-Eval") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("DataXAI/AnesTRACE-Eval") model = AutoModelForMultimodalLM.from_pretrained("DataXAI/AnesTRACE-Eval", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use DataXAI/AnesTRACE-Eval with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "DataXAI/AnesTRACE-Eval" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "DataXAI/AnesTRACE-Eval", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/DataXAI/AnesTRACE-Eval
- SGLang
How to use DataXAI/AnesTRACE-Eval with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "DataXAI/AnesTRACE-Eval" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "DataXAI/AnesTRACE-Eval", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "DataXAI/AnesTRACE-Eval" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "DataXAI/AnesTRACE-Eval", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use DataXAI/AnesTRACE-Eval with Docker Model Runner:
docker model run hf.co/DataXAI/AnesTRACE-Eval
AnesTRACE-Eval
AnesTRACE-Eval is a specialized evaluator for AnesTRACE, a benchmark of clinical reasoning and sequential decision-making in anesthesia and perioperative care. The model assigns structured scores to candidate responses using the AnesTRACE evaluation rubrics.
This repository contains the merged inference checkpoint. It can be loaded directly with Transformers and does not require a separate LoRA adapter.
Intended use
AnesTRACE-Eval is designed for research evaluation of model outputs on AnesTRACE tasks, including:
- Level Two single-point perioperative decision-making;
- Level Three multi-turn clinical reasoning and intervention decisions;
- clinical correctness;
- evidence-based reasoning and grounding;
- task completeness;
- safety severity for intervention-related outputs;
- temporal adaptation and longitudinal management coherence for multi-turn trajectories.
The model is intended to be used with the official AnesTRACE evaluator prompts and output schemas. It is not intended to generate or validate autonomous clinical care.
Model details
| Item | Value |
|---|---|
| Model name | AnesTRACE-Eval |
| Base model | Qwen/Qwen3.5-9B |
| Architecture | Qwen3.5 conditional generation model |
| Primary language | English |
| Parameter precision | BF16 |
| Context length in configuration | 262,144 tokens |
| Training framework | LLaMA-Factory |
| Final alignment method | Direct Preference Optimization (DPO) with LoRA |
| Release format | Merged safetensors checkpoint |
The vision tower was frozen during fine-tuning. The released evaluator is used as a text-based judge in the AnesTRACE evaluation pipeline.
Training
Training was performed in two stages.
Stage 1: supervised evaluator fine-tuning
The base Qwen3.5-9B model was fine-tuned on AnesTRACE evaluator examples using LoRA. The supervised data teach the model the Level Two and Level Three evaluation rubrics, structured score schemas, and safety labels.
Key SFT settings:
| Setting | Value |
|---|---|
| LoRA rank | 16 |
| LoRA alpha | 32 |
| LoRA dropout | 0.05 |
| Learning rate | 1e-4 |
| Epochs | 5 |
| Sequence cutoff | 8,192 tokens |
| Scheduler | Cosine |
| Warmup ratio | 0.05 |
| Precision | BF16 |
Stage 2: preference optimization
The merged SFT model was further aligned using true preference pairs derived from Level Three action evaluation data. DPO was applied through a new LoRA adapter, which was subsequently merged into the SFT model.
Key DPO settings:
| Setting | Value |
|---|---|
| Objective | Sigmoid DPO |
| Preference beta | 0.1 |
| LoRA rank | 8 |
| LoRA alpha | 16 |
| LoRA dropout | 0.05 |
| Learning rate | 3e-6 |
| Epochs | 1 |
| Sequence cutoff | 6,144 tokens |
| Per-device training batch size | 1 |
| Gradient accumulation steps | 16 |
| Scheduler | Cosine |
| Warmup ratio | 0.05 |
| Seed | 42 |
| Precision | BF16 |
| Released checkpoint | Step 105 |
The final checkpoint contains the merged model weights in four safetensors shards.
Usage
Install a recent Transformers version with Qwen3.5 support:
pip install -U "transformers>=5.8.0" accelerate safetensors
The following example performs deterministic text-only evaluation. Replace the abbreviated prompts with the official AnesTRACE system prompt and case input for the target evaluation level.
import torch
from transformers import AutoProcessor, Qwen3_5ForConditionalGeneration
model_id = "DataXAI/AnesTRACE-Eval"
processor = AutoProcessor.from_pretrained(
model_id,
trust_remote_code=True,
)
model = Qwen3_5ForConditionalGeneration.from_pretrained(
model_id,
torch_dtype=torch.bfloat16,
device_map="auto",
trust_remote_code=True,
)
messages = [
{
"role": "system",
"content": "You are the AnesTRACE clinical evaluation model. Follow the supplied rubric and return only the required JSON object.",
},
{
"role": "user",
"content": "[Evaluation Instruction]\nEvaluate the candidate response using the official AnesTRACE rubric.\n\n[Case]\n...\n\n[Reference Answer]\n...\n\n[Candidate Answer]\n...",
},
]
inputs = processor.apply_chat_template(
messages,
tokenize=True,
add_generation_prompt=True,
return_dict=True,
return_tensors="pt",
enable_thinking=False,
)
inputs = inputs.to(model.device)
with torch.inference_mode():
generated = model.generate(
**inputs,
max_new_tokens=2048,
do_sample=False,
)
new_tokens = generated[:, inputs["input_ids"].shape[1]:]
output = processor.batch_decode(
new_tokens,
skip_special_tokens=True,
clean_up_tokenization_spaces=False,
)[0]
print(output)
For reproducible benchmark scoring, use the complete official prompt for the selected level, disable sampling, validate the returned JSON against the corresponding schema, and derive aggregate totals from the validated dimension scores.
Output interpretation
AnesTRACE-Eval uses discrete rubric scores. For Level Two, each B1–B4 task is evaluated independently on:
d1_clinical_correctness;d2_evidence_based_reasoning;d3_task_completeness.
Each dimension uses an integer score from 0 to 2. Intervention Decision and Reassessment Plan additionally receive a safety severity label: safe, minor, major, or critical.
Level Three turn evaluation applies the same three dimensions to diagnosis and intervention decisions. A separate trajectory evaluation measures temporal evidence and response adaptation, and longitudinal management coherence.
The evaluator's raw output should be schema-validated before scores are aggregated. Application code should recompute total scores from validated dimension scores rather than trusting a model-generated total.
Limitations
- AnesTRACE-Eval is a learned evaluator and can make scoring or calibration errors.
- Scores may be sensitive to prompt formatting, missing reference evidence, truncated candidate answers, and outputs that do not follow the expected schema.
- The model was optimized for the English AnesTRACE evaluation format. Performance on other languages, unrelated medical specialties, or arbitrary free-form judging tasks has not been established.
- Agreement with this evaluator does not establish clinical correctness or patient safety.
- The model must not be used as a medical device, for diagnosis or treatment, or as the sole basis for clinical, regulatory, or deployment decisions.
- High-stakes results should be reviewed by qualified clinicians, with disagreement analysis and human adjudication where appropriate.
Data and privacy
The model is intended for evaluation on de-identified research data. Users are responsible for ensuring that inputs comply with applicable privacy, institutional, and data-governance requirements. Do not submit identifiable patient information.
License
This model is released under the Apache 2.0 license, subject to the terms and restrictions of the Qwen3.5 base model and any applicable AnesTRACE data licenses.
Citation
If you use AnesTRACE-Eval, please cite our work.
@misc{huang2026anestrace,
title={AnesTRACE: Benchmarking Intraoperative Anesthesia from Multimodal Perception to Multi-step Decision-Making},
author={Huang, Ziwei and Gao, Qi and Ji, Zhe and Yao, Yuanyuan and Zhang, Fengjiang and Yan, Min and Xie, Zhongle and Chen, Gang},
year={2026},
eprint={2609.32740},
archivePrefix={arXiv},
url={https://arxiv.org/abs/2609.32740}
}
Acknowledgements
AnesTRACE-Eval was developed using Qwen3.5 and LLaMA-Factory. We thank the developers and maintainers of these projects.
- Downloads last month
- 174