YAML Metadata Warning:empty or missing yaml metadata in repo card

Check out the documentation for more information.

Taiwan Medical Named Entity Recognition (NER) Model

This directory contains a fine-tuned BERT-based Token Classification model tailored for Taiwanese medical terms, surgical procedures, and medical devices.

๐Ÿท๏ธ Model Labels

The model identifies the following three entity classes:

  1. BODY_PART: Anatomical regions, organs, or physiological systems (e.g. ๅไบŒๆŒ‡่…ธ, ไนณๆˆฟ, ่ง’่†œ, ่…Ž่‡Ÿ).
  2. SURGERY_TYPE: Surgical actions, techniques, or procedure types (e.g. ๅˆ‡้™ค่ก“, ็ธซๅˆ่ก“, ็งปๆค, ็ฝฎๆ›่ก“).
  3. TECH_DEVICE: Medical devices, surgical technologies, or specific equipment (e.g. ้”ๆ–‡่ฅฟ, ้›ทๅฐ„, ่…น่…”้ก, ่ถ…้ŸณๆณขไนณๅŒ–).

๐Ÿ“‚ Included Files

  • model.safetensors: Fine-tuned weights.
  • config.json: Architecture, hyperparameters, and label definitions (id2label / label2id).
  • tokenizer.json: Serialized tokenizer configuration.
  • tokenizer_config.json: Instantiate arguments.

๐Ÿš€ Quick Start Sample Code

Below is a complete, self-contained Python script to load the model and run inference:

import os
import torch
from transformers import AutoTokenizer, AutoModelForTokenClassification

# Model folder (local path containing these files)
MODEL_DIR = os.path.dirname(os.path.abspath(__file__))

LABEL_LIST = ["O", "B-BODY_PART", "I-BODY_PART", "B-SURGERY_TYPE", "I-SURGERY_TYPE", "B-TECH_DEVICE", "I-TECH_DEVICE"]

def extract_entities(text: str, tokenizer, model) -> list:
    """Tokenizes text and groups token classifications into character-aligned entity spans."""
    if not text.strip():
        return []
        
    inputs = tokenizer(
        text, 
        return_offsets_mapping=True, 
        return_tensors="pt",
        truncation=True,
        max_length=512
    )
    
    device = next(model.parameters()).device
    inputs = {k: v.to(device) for k, v in inputs.items()}
    
    with torch.no_grad():
        outputs = model(**{k: v for k, v in inputs.items() if k != "offset_mapping"})
        
    logits = outputs.logits
    predictions = torch.argmax(logits, dim=2)[0].cpu().numpy()
    offsets = inputs["offset_mapping"][0].cpu().numpy()
    
    entities = []
    current_entity = None
    
    for idx, offset in enumerate(offsets):
        start, end = offset
        # Ignore padding and special tokens
        if start == 0 and end == 0:
            continue
            
        label = LABEL_LIST[predictions[idx]]
        
        if label.startswith("B-"):
            if current_entity:
                entities.append(current_entity)
            entity_type = label.split("-")[1]
            current_entity = {
                "label": entity_type,
                "start": int(start),
                "end": int(end),
                "text": text[start:end]
            }
        elif label.startswith("I-"):
            entity_type = label.split("-")[1]
            if current_entity and current_entity["label"] == entity_type:
                current_entity["end"] = int(end)
                current_entity["text"] = text[current_entity["start"]:int(end)]
            else:
                if current_entity:
                    entities.append(current_entity)
                current_entity = {
                    "label": entity_type,
                    "start": int(start),
                    "end": int(end),
                    "text": text[start:end]
                }
        else: # 'O'
            if current_entity:
                entities.append(current_entity)
                current_entity = None
                
    if current_entity:
        entities.append(current_entity)
        
    return entities

def main():
    print(f"Loading model from {MODEL_DIR}...")
    tokenizer = AutoTokenizer.from_pretrained(MODEL_DIR)
    model = AutoModelForTokenClassification.from_pretrained(MODEL_DIR)
    
    # Detect Apple Silicon GPU (mps), CUDA, or CPU
    device = "cpu"
    if torch.backends.mps.is_available():
        device = "mps"
    elif torch.cuda.is_available():
        device = "cuda"
    model = model.to(device)
    model.eval()
    
    # Example Inference
    test_sentence = "็—…ๆ‚ฃๅ› ๅณๅด่…น่‚กๆบ็–ๆฐฃไฝ้™ข๏ผŒๅœจ้–€่จบๅฏฆๆ–ฝไบ†ๅพฎๅ‰ตๅ…ง่ฆ–้ก็–ๆฐฃไฟฎ่ฃœ่ก“ใ€‚"
    entities = extract_entities(test_sentence, tokenizer, model)
    
    print(f"\nInput: {test_sentence}")
    print("Extracted Entities:")
    for ent in entities:
        print(f"  - {ent['text']} | Label: {ent['label']} | Spans: ({ent['start']}, {ent['end']})")

if __name__ == "__main__":
    main()

๐Ÿ› ๏ธ Requirements

Make sure you have PyTorch and Hugging Face Transformers installed:

pip install torch transformers
Downloads last month
2
Safetensors
Model size
0.1B params
Tensor type
F32
ยท
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support