DT4H_CardioBERTa_parents_sv_only_snomed

DT4H_CardioBERTa_parents_sv_only_snomed is a Swedish biomedical terminology encoder for clinical concept normalization and entity linking. It is initialized from [DT4H/CardioBERTa.sv] and specialized using CUI-supervised terminology pairs and metric learning.

Backbone

The backbone belongs to the CardioBERTa family from CardioLM - a multilingual suite of small language models for the cardiology domain. CardioBERTa comprises language-specific encoder models adapted to cardiology through continued pretraining on monolingual biomedical and cardiology-related corpora using Masked Language Modeling (MLM). The family covers Czech, Dutch, English, Italian, Romanian, Spanish and Swedish.

Training

Language Swedish (sv)
Triplet collection only_snomed
Strategy parents
Objective Multi-Similarity Loss
Mining All triplets, margin 0.2
Pooling CLS
Epochs 1
Batch size 256
Learning rate 2e-5
Max. length 25

CUI-supervised terminology pairs enriched with parent-level ontology relations.

Terminology statistics

Strategy Triplets CUIs Unique terms Unique positives Terms/CUI Δ terms
synonyms 71,919 71,919 141,369 71,328 2.00 0
parents 1,024,243 398,217 389,241 280,684 3.57 +247,872
grandparents 3,227,930 398,450 389,241 323,085 9.10 +247,872

This model uses 1,024,243 triplets, covering 398,217 CUIs and 389,241 unique normalized terms.

The training terminology is not distributed with this repository because it contains resources subject to UMLS licensing conditions. Only aggregate statistics are released.

Intended use

The model is intended for terminology embedding, biomedical candidate retrieval, concept normalization and entity linking, particularly in cardiology and clinical NLP pipelines. It is not intended for direct clinical decision-making.

Usage

import torch
import torch.nn.functional as F
from transformers import AutoModel, AutoTokenizer

model_id = "DT4H/DT4H_CardioBERTa_parents_sv_only_snomed"

tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModel.from_pretrained(model_id)

inputs = tokenizer(
    "clinical concept",
    return_tensors="pt",
    truncation=True,
    max_length=25,
)

with torch.no_grad():
    output = model(**inputs)

embedding = F.normalize(
    output.last_hidden_state[:, 0, :],
    p=2,
    dim=1,
)

Reference

Danu et al. CardioLM - a multilingual suite of small language models for the cardiology domain.

Developed within the DataTools4Heart (DT4H) project, Grant Agreement 101057849.

Downloads last month
10
Safetensors
Model size
0.1B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for DT4H/CardioBERTa.sv_P_only_snomed

Finetuned
(12)
this model

Collection including DT4H/CardioBERTa.sv_P_only_snomed