Instructions to use DT4H/CardioBERTa.it_translations_only with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use DT4H/CardioBERTa.it_translations_only with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("feature-extraction", model="DT4H/CardioBERTa.it_translations_only")# Load model directly from transformers import AutoTokenizer, AutoModel tokenizer = AutoTokenizer.from_pretrained("DT4H/CardioBERTa.it_translations_only") model = AutoModel.from_pretrained("DT4H/CardioBERTa.it_translations_only", device_map="auto") - Notebooks
- Google Colab
- Kaggle
DT4H_CardioBERTa_it_translations_only
DT4H_CardioBERTa_it_translations_only is a Italian biomedical terminology encoder for clinical concept normalization and entity linking. It is initialized from [DT4H/CardioBERTa.it] and specialized using CUI-supervised terminology pairs and metric learning.
Backbone
The backbone belongs to the CardioBERTa family from CardioLM - a multilingual suite of small language models for the cardiology domain. CardioBERTa comprises language-specific encoder models adapted to cardiology through continued pretraining on monolingual biomedical and cardiology-related corpora using Masked Language Modeling (MLM). The family covers Czech, Dutch, English, Italian, Romanian, Spanish and Swedish.
Training
| Language | Italian (it) |
| Triplet collection | translations_only |
| Strategy | synonyms |
| Objective | Multi-Similarity Loss |
| Mining | All triplets, margin 0.2 |
| Pooling | CLS |
| Epochs | 1 |
| Batch size | 256 |
| Learning rate | 2e-5 |
| Max. length | 25 |
CUI-supervised synonym pairs.
Terminology statistics
| Strategy | Triplets | CUIs | Unique terms | Unique positives | Terms/CUI | Δ terms |
|---|---|---|---|---|---|---|
| synonyms | 69,631 | 69,631 | 136,720 | 69,113 | 2.00 | 0 |
| parents | 1,597,673 | 476,349 | 529,199 | 414,377 | 3.93 | +392,479 |
| grandparents | 4,714,271 | 476,970 | 529,487 | 461,834 | 9.83 | +392,767 |
This model uses 69,631 triplets, covering 69,631 CUIs and 136,720 unique normalized terms.
The training terminology is not distributed with this repository because it contains resources subject to UMLS licensing conditions. Only aggregate statistics are released.
Intended use
The model is intended for terminology embedding, biomedical candidate retrieval, concept normalization and entity linking, particularly in cardiology and clinical NLP pipelines. It is not intended for direct clinical decision-making.
Usage
import torch
import torch.nn.functional as F
from transformers import AutoModel, AutoTokenizer
model_id = "DT4H/DT4H_CardioBERTa_it_translations_only"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModel.from_pretrained(model_id)
inputs = tokenizer(
"clinical concept",
return_tensors="pt",
truncation=True,
max_length=25,
)
with torch.no_grad():
output = model(**inputs)
embedding = F.normalize(
output.last_hidden_state[:, 0, :],
p=2,
dim=1,
)
Reference
Danu et al. CardioLM - a multilingual suite of small language models for the cardiology domain.
Developed within the DataTools4Heart (DT4H) project, Grant Agreement 101057849.
- Downloads last month
- 7
Model tree for DT4H/CardioBERTa.it_translations_only
Base model
DT4H/CardioBERTa.it