Instructions to use TCMLLM/Lingdan-8B-SFT with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use TCMLLM/Lingdan-8B-SFT with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="TCMLLM/Lingdan-8B-SFT") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("TCMLLM/Lingdan-8B-SFT") model = AutoModelForCausalLM.from_pretrained("TCMLLM/Lingdan-8B-SFT", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use TCMLLM/Lingdan-8B-SFT with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "TCMLLM/Lingdan-8B-SFT" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "TCMLLM/Lingdan-8B-SFT", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/TCMLLM/Lingdan-8B-SFT
- SGLang
How to use TCMLLM/Lingdan-8B-SFT with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "TCMLLM/Lingdan-8B-SFT" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "TCMLLM/Lingdan-8B-SFT", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "TCMLLM/Lingdan-8B-SFT" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "TCMLLM/Lingdan-8B-SFT", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use TCMLLM/Lingdan-8B-SFT with Docker Model Runner:
docker model run hf.co/TCMLLM/Lingdan-8B-SFT
Lingdan-8B-SFT
Lingdan-V2: Large Language Models for Clinically Grounded Reasoning in Traditional Chinese Medicine
Model · Results · Quickstart · vLLM · Prompts · Citation
Model Details
Lingdan-8B-SFT is a Qwen3-based TCM model in the Lingdan-V2 family. CPT → SFT on 7,729,982 samples (2.08B tokens).
| Field | Value |
|---|---|
| Parameter scale | 8B |
| Architecture family | Qwen3 causal language model |
| Training stage | Supervised fine-tuning (SFT) |
| Starting checkpoint / backbone | Lingdan-8B-Base |
| Intended use | TCM instruction following, question answering, information extraction, and diagnostic-therapeutic tasks |
| Platform model ID | TCMLLM/Lingdan-8B-SFT |
Checkpoint Evaluation
Paper-reported results for Lingdan-8B-SFT. Full: LingLan (25,620 items); Hard: LingLan-Hard (5,200 items). Overall Score averages 16 primary metrics, excluding dose MAE. All scores use a 0–100 scale except MAE (lower is better). The paper uses zero-shot Chinese prompts, temperature 0.6, and at most 8,192 generated tokens; reproduction also requires the same data and scoring pipeline.
| Metric | LingLan Full | LingLan-Hard |
|---|---|---|
| Overall Score ↑ | 63.0 | 36.6 |
| Prescription generation F1 ↑ | 29.8 | 16.1 |
| Private prescription evaluation (4,348 cases) | Score |
|---|---|
| Herb F1 ↑ | 34.2 |
| Herb precision ↑ | 52.7 |
| Herb recall ↑ | 26.0 |
| Dose cosine ↑ | 34.2 |
| Dose MAE ↓ | 5.812 |
The private set contains 4,348 de-identified spleen/stomach disorder cases. See the full comparison for other checkpoints.
Quickstart
Use Transformers ≥4.51.0; see the Qwen3 quickstart.
pip install -U torch transformers accelerate
from transformers import AutoModelForCausalLM, AutoTokenizer
model_name = "TCMLLM/Lingdan-8B-SFT"
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForCausalLM.from_pretrained(
model_name,
torch_dtype="auto",
device_map="auto",
).eval()
messages = [{"role": "user", "content": "请简要介绍中医辨证论治的基本概念。"}]
text = tokenizer.apply_chat_template(
messages,
tokenize=False,
add_generation_prompt=True,
enable_thinking=True,
)
inputs = tokenizer([text], return_tensors="pt").to(model.device)
outputs = model.generate(
**inputs,
max_new_tokens=8192,
eos_token_id=[151645, 151643], # <|im_end|> and <|endoftext|>
do_sample=True,
temperature=0.6,
top_p=0.95,
top_k=20,
)
new_tokens = outputs[0, inputs.input_ids.shape[1]:]
print(tokenizer.decode(new_tokens, skip_special_tokens=True))
The explicit EOS IDs stop generation at either the chat-turn or text terminator.
Thinking Mode
The quickstart enables thinking with enable_thinking=True. Set it to False for direct answers. The released chat templates support this switch.
| Mode | enable_thinking |
Temperature | Top-p | Top-k |
|---|---|---|---|---|
| Thinking | True |
0.6 | 0.95 | 20 |
| Non-thinking | False |
0.7 | 0.8 | 20 |
These are Qwen3 inference starting settings; behavior after fine-tuning may vary. Prefer the explicit switch over the unverified /think and /no_think soft instructions. Separate direct-generation output as follows:
raw_output = tokenizer.decode(new_tokens, skip_special_tokens=True)
reasoning_part, delimiter, final_answer = raw_output.partition("</think>")
if delimiter:
thinking_content = reasoning_part.removeprefix("<think>").strip()
final_answer = final_answer.strip()
else:
thinking_content = ""
final_answer = raw_output.strip()
print("Thinking:", thinking_content)
print("Answer:", final_answer)
If generation hits the token limit before the thinking block closes, increase the budget and retry; the answer is incomplete.
vLLM Deployment
Install vLLM in a GPU environment with compatible CUDA/PyTorch; see the Qwen3 vLLM guide.
pip install -U "vllm>=0.9.0" openai
Start the Server
Hugging Face:
export VLLM_USE_MODELSCOPE=false
vllm serve TCMLLM/Lingdan-8B-SFT \
--served-model-name Lingdan-8B-SFT \
--host 127.0.0.1 \
--port 8000 \
--max-model-len 16384 \
--reasoning-parser qwen3
Replace the repository ID with a local model directory to reuse downloaded files. For two GPUs, append --tensor-parallel-size 2. Adjust the 16,384-token context limit (input + output) to available GPU memory.
Chat Client
from openai import OpenAI
client = OpenAI(base_url="http://127.0.0.1:8000/v1", api_key="EMPTY")
response = client.chat.completions.create(
model="Lingdan-8B-SFT",
messages=[{"role": "user", "content": "请简要介绍中医辨证论治的基本概念。"}],
temperature=0.6,
top_p=0.95,
max_tokens=8192,
extra_body={
"top_k": 20,
"chat_template_kwargs": {"enable_thinking": True},
"stop_token_ids": [151645, 151643],
},
)
message = response.choices[0].message
thinking_content = (
getattr(message, "reasoning", None)
or getattr(message, "reasoning_content", None)
or ""
)
print("Thinking:", thinking_content)
print("Answer:", message.content)
print("Finish reason:", response.choices[0].finish_reason)
chat_template_kwargs selects thinking mode; --reasoning-parser qwen3 separates its output. To disable thinking, use the non-thinking settings above. The client supports both reasoning and the older reasoning_content field (vLLM reference). Parse task JSON from message.content; finish_reason="length" indicates truncation.
EMPTY is a placeholder for the local server started without API-key authentication.
Task Prompts
For SFT/R1 models, fill the placeholders in these LingLan evaluation templates and pass the text as the user message.
Clinical-record Entity Extraction
Source: LingLan ner_clinic.py.
请从以下病历文本中抽取指定类型的实体。
注意:
1. 请先仔细分析文本内容,思考每个实体类型的识别要点,然后再输出结果
2. 如果某个实体类型在文本中没有对应实体,请返回空列表[]
3. 如果同一个实体在文本中出现多次,请在列表中按顺序重复出现
文本:
{text}
需要抽取的实体类型:
{entity_types}
请按照以下JSON格式输出结果,每个实体类型对应一个列表,如果文本中某个实体出现多次,请在列表中重复出现:
```json
{
"实体类型1": ["实体1", "实体2", ...],
"实体类型2": ["实体1", "实体2", ...],
...
}
```
Classical-text Entity Extraction
Source: LingLan ner_ancient.py.
请从以下古籍文本中抽取指定类型的实体。
注意:
1. 请先仔细分析古籍文本内容,思考古文语言特点和各实体类型的识别要点,然后再输出结果
2. 如果某个实体类型在文本中没有对应实体,请返回空列表[]
3. 如果同一个实体在文本中出现多次,请在列表中按顺序重复出现
文本:
{text}
需要抽取的实体类型:
{entity_types}
请按照以下JSON格式输出结果,每个实体类型对应一个列表,如果文本中某个实体出现多次,请在列表中重复出现:
```json
{
"实体类型1": ["实体1", "实体2", ...],
"实体类型2": ["实体1", "实体2", ...],
...
}
```
Joint Syndrome, Treatment, and Prescription Generation
Source: LingLan pres_diag_eval.py.
请根据以下病历文本,进行完整的中医诊疗分析,包括辨证、治法和处方。
注意:
1. 请仔细分析患者的症状表现、舌脉象等信息
2. 辨证:给出准确的中医证型判断(如有多个,用逗号分隔)
3. 治法:根据辨证结果确定相应的治疗方法(如有多个,用逗号分隔)
4. 处方:开具具体的中药处方,格式为"药名: 剂量, 药名: 剂量"
病历文本:
{medical_record}
请按照以下JSON格式输出结果:
```json
{
"辨证": "证型1, 证型2, ...",
"治法": "治法1, 治法2, ...",
"处方": "药名1: 剂量1, 药名2: 剂量2, 药名3: 剂量3, ..."
}
```
Use the dataset’s exact entity labels. JSON examples contain illustrative ellipses; actual output must be valid JSON. In joint generation, 辨证, 治法, and 处方 are strings; format prescription items as 药名: 数值g. Extract task JSON from the final answer using the output handling above.
Data Privacy and Responsible Use
Lingdan-V2 is for research and does not replace qualified clinical judgment. Private clinical data are not released. Retrospective prescription agreement is not evidence of clinical safety or effectiveness; use de-identified inputs and professional review for any clinical application.
License
Lingdan-V2 model weights and project-owned documentation and images are released under the Apache License 2.0. See NOTICE for attribution notices. Third-party data and materials remain subject to their original terms.
Citation
Cite Lingdan-V2 and the earlier Lingdan work:
@misc{hua2026lingdanv2,
title = {{Lingdan-V2}: Large language models for clinically grounded reasoning in traditional Chinese medicine},
author = {Hua, Rui and Wei, Yu and Jin, Ziri and Sun, Zhuo and Chang, Kai and Xia, Jianan and Shu, Zixin and Li, Xiaodong and Jia, Dongmei and Dong, Fei and Zhang, Runshun and Yu, Jian and Xu, Hao and Yu, Haibin and Li, Jiansheng and Liu, Baoyan and Wang, Wenjia and Zhou, Xuezhong},
year = {2026},
url = {https://github.com/TCMAI-BJTU/Lingdan-V2}
}
@article{hua2024lingdan,
title={Lingdan: enhancing encoding of traditional Chinese medicine knowledge for clinical reasoning tasks with large language models},
author={Hua, Rui and Dong, Xin and Wei, Yu and Shu, Zixin and Yang, Pengcheng and Hu, Yunhui and Zhou, Shuiping and Sun, He and Yan, Kaijing and Yan, Xijun and others},
journal={Journal of the American Medical Informatics Association},
volume={31},
number={9},
pages={2019--2029},
year={2024},
publisher={Oxford University Press}
}
Related Links
GitHub · Hugging Face collection · ModelScope collection · LingLan benchmark
- Downloads last month
- 25