Instructions to use Aditya109/us-openfda-drug-embed-v1 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- sentence-transformers
How to use Aditya109/us-openfda-drug-embed-v1 with sentence-transformers:
from sentence_transformers import SentenceTransformer model = SentenceTransformer("Aditya109/us-openfda-drug-embed-v1") sentences = [ "Brand: cough drops dollar general menthol 80ct. Ingredients: menthol. Strength: 5.4 mg/1. Form: lozenge. Route: oral.", "Brand: dermfree numbing. Ingredients: menthol. Strength: 5 g/100g. Form: cream. Route: topical.", "Brand: equate menthol cough drops. Ingredients: menthol. Strength: 5.4 mg/1. Form: lozenge. Route: oral.", "Brand: medline. Ingredients: cetirizine hydrochloride. Strength: 10 mg/1. Form: tablet. Route: oral." ] embeddings = model.encode(sentences) similarities = model.similarity(embeddings, embeddings) print(similarities.shape) # [4, 4] - Notebooks
- Google Colab
- Kaggle
SentenceTransformer based on FremyCompany/BioLORD-2023
U.S. openFDA Drug Product Research Model
This experimental model was fine-tuned on structured drug-product records from the official openFDA National Drug Code Directory. Training text includes brand name, active ingredients, strength, dosage form, and route.
The openFDA source data is generally available under CC0/public-domain terms; the base model and software dependencies retain their own terms. This model is for research and catalog-matching experiments. It must not be used for diagnosis, prescribing, dosage decisions, interaction checks, or automatic medicine substitution.
Held-out evaluation:
- nearest-neighbour exact-product accuracy: 0.6585 -> 0.8171
- validation triplet accuracy: 0.8995 -> 0.9416
- normalized separation margin: 2.5155 -> 2.8190
This is a sentence-transformers model finetuned from FremyCompany/BioLORD-2023 on the pairs and triplets datasets. It maps sentences & paragraphs to a 768-dimensional dense vector space and can be used for semantic textual similarity, semantic search, paraphrase mining, classification, clustering, and more.
Model Details
Model Description
- Model Type: Sentence Transformer
- Base model: FremyCompany/BioLORD-2023
- Maximum Sequence Length: 32 tokens
- Output Dimensionality: 768 dimensions
- Similarity Function: Cosine Similarity
- Supported Modality: Text
- Training Datasets:
- pairs
- triplets
Model Sources
- Documentation: Sentence Transformers Documentation
- Repository: Sentence Transformers on GitHub
- Hugging Face: Sentence Transformers on Hugging Face
Full Model Architecture
SentenceTransformer(
(0): Transformer({'transformer_task': 'feature-extraction', 'modality_config': {'text': {'method': 'forward', 'method_output_name': 'last_hidden_state'}}, 'module_output_name': 'token_embeddings', 'architecture': 'MPNetModel'})
(1): Pooling({'embedding_dimension': 768, 'pooling_mode': 'mean', 'include_prompt': True})
)
Usage
Direct Usage (Sentence Transformers)
First install the Sentence Transformers library:
pip install -U sentence-transformers
Then you can load this model and run inference.
from sentence_transformers import SentenceTransformer
# Download from the 🤗 Hub
model = SentenceTransformer("sentence_transformers_model_id")
# Run inference
sentences = [
'Brand: smart care island mango hand sanitizer. Ingredients: alcohol. Strength: 70 ml/100ml. Form: spray. Route: topical.',
'Brand: ashleybelle moisturizing hand sanitizer cranberry sage. Ingredients: alcohol. Strength: 70 ml/100ml. Form: spray. Route: topical.',
'Brand: walgreens dandruff itchy dry scalp defense anti-dandruff. Ingredients: pyrithione zinc. Strength: 10 mg/ml. Form: shampoo. Route: topical.',
]
embeddings = model.encode(sentences)
print(embeddings.shape)
# [3, 768]
# Get the similarity scores for the embeddings
similarities = model.similarity(embeddings, embeddings)
print(similarities)
# tensor([[1.0000, 0.8526, 0.4529],
# [0.8526, 1.0000, 0.4206],
# [0.4529, 0.4206, 1.0000]])
Evaluation
Metrics
Composition
- Dataset:
comp - Evaluated with
main.CompositionEvaluator
| Metric | Value |
|---|---|
| same | 0.9165 |
| diff | 0.5142 |
| margin | 0.4023 |
| margin_norm | 2.819 |
| nn_accuracy | 0.8171 |
Triplet
- Dataset:
val - Evaluated with
TripletEvaluator
| Metric | Value |
|---|---|
| cosine_accuracy | 0.9416 |
Training Details
Training Datasets
pairs
- Dataset: pairs
- Size: 13,331 training samples
- Columns:
anchorandpositive - Approximate statistics based on the first 100 samples:
anchor positive type string string modality text text details - min: 27 tokens
- mean: 31.86 tokens
- max: 32 tokens
- min: 28 tokens
- mean: 31.86 tokens
- max: 32 tokens
- Samples:
anchor positive Brand: gps topical anesthetic. Ingredients: benzocaine. Strength: 200 mg/g. Form: gel. Route: dental.Brand: primo topical anesthetic. Ingredients: benzocaine. Strength: 200 mg/g. Form: gel. Route: dental.Brand: bencocaine topical anesthetic. Ingredients: benzocaine. Strength: 200 mg/g. Form: gel. Route: dental.Brand: tiger supply inc topical anesthetic. Ingredients: benzocaine. Strength: 200 mg/g. Form: gel. Route: dental.Brand: safco sensicaine ultra topical anesthetic gel. Ingredients: benzocaine. Strength: 200 mg/g. Form: gel. Route: dental.Brand: advance topical anesthetic gel. Ingredients: benzocaine. Strength: 200 mg/g. Form: gel. Route: dental. - Loss:
CachedGISTEmbedLosswith these parameters:{ "guide": "SentenceTransformer(None)", "temperature": 0.01, "mini_batch_size": 64, "mini_batch_num_tokens": null, "margin_strategy": "absolute", "margin": 0.0, "contrast_anchors": true, "contrast_positives": true, "gather_across_devices": false }
triplets
- Dataset: triplets
- Size: 27,325 training samples
- Columns:
anchor,positive, andnegative - Approximate statistics based on the first 100 samples:
anchor positive negative type string string string modality text text text details - min: 27 tokens
- mean: 31.8 tokens
- max: 32 tokens
- min: 28 tokens
- mean: 31.88 tokens
- max: 32 tokens
- min: 28 tokens
- mean: 31.61 tokens
- max: 32 tokens
- Samples:
anchor positive negative Brand: candee caine topical anesthetic. Ingredients: benzocaine. Strength: 200 mg/g. Form: gel. Route: dental.Brand: advance topical anesthetic gel. Ingredients: benzocaine. Strength: 200 mg/g. Form: gel. Route: dental.Brand: 7 select oral pain maximum strength relief. Ingredients: benzocaine. Strength: 200 mg/g. Form: gel. Route: topical.Brand: kolorz topical anesthetic blue rasberry. Ingredients: benzocaine. Strength: 200 mg/g. Form: gel. Route: dental.Brand: kolorz topical anesthetic triple mint. Ingredients: benzocaine. Strength: 200 mg/g. Form: gel. Route: dental.Brand: lubelife climax control delay. Ingredients: benzocaine. Strength: 7.5 g/100ml. Form: spray. Route: topical.Brand: quala topical anesthetic gel. Ingredients: benzocaine. Strength: 200 mg/g. Form: gel. Route: dental.Brand: kolorz topical anesthetic cotton candy. Ingredients: benzocaine. Strength: 200 mg/g. Form: gel. Route: dental.Brand: hurricaine topical anesthetic. Ingredients: benzocaine. Strength: 200 mg/g. Form: gel. Route: dental | periodontal. - Loss:
CachedGISTEmbedLosswith these parameters:{ "guide": "SentenceTransformer(None)", "temperature": 0.01, "mini_batch_size": 64, "mini_batch_num_tokens": null, "margin_strategy": "absolute", "margin": 0.0, "contrast_anchors": true, "contrast_positives": true, "gather_across_devices": false }
Training Hyperparameters
Non-Default Hyperparameters
per_device_train_batch_size: 512num_train_epochs: 1.0learning_rate: 1e-05warmup_steps: 0.1bf16: Trueload_best_model_at_end: Truebatch_sampler: no_duplicates
All Hyperparameters
Click to expand
per_device_train_batch_size: 512num_train_epochs: 1.0max_steps: -1learning_rate: 1e-05lr_scheduler_type: linearlr_scheduler_kwargs: Nonewarmup_steps: 0.1optim: adamw_torch_fusedoptim_args: Noneweight_decay: 0.0adam_beta1: 0.9adam_beta2: 0.999adam_epsilon: 1e-08optim_target_modules: Nonegradient_accumulation_steps: 1average_tokens_across_devices: Truemax_grad_norm: 1.0label_smoothing_factor: 0.0bf16: Truefp16: Falsebf16_full_eval: Falsefp16_full_eval: Falsetf32: Nonegradient_checkpointing: Falsegradient_checkpointing_kwargs: Nonetorch_compile: Falsetorch_compile_backend: Nonetorch_compile_mode: Noneuse_liger_kernel: Falseliger_kernel_config: Noneuse_cache: Falseneftune_noise_alpha: Nonetorch_empty_cache_steps: Noneauto_find_batch_size: Falselog_on_each_node: Truelogging_nan_inf_filter: Trueinclude_num_input_tokens_seen: nolog_level: passivelog_level_replica: warningdisable_tqdm: Falseproject: huggingfacetrackio_space_id: Nonetrackio_bucket_id: Nonetrackio_static_space_id: Noneper_device_eval_batch_size: 8prediction_loss_only: Trueeval_on_start: Falseeval_do_concat_batches: Trueeval_use_gather_object: Falseeval_accumulation_steps: Noneinclude_for_metrics: []batch_eval_metrics: Falsesave_only_model: Falsesave_on_each_node: Falseenable_jit_checkpoint: Falsepush_to_hub: Falsehub_private_repo: Nonehub_model_id: Nonehub_strategy: every_savehub_always_push: Falsehub_revision: Noneload_best_model_at_end: Trueignore_data_skip: Falserestore_callback_states_from_checkpoint: Falsefull_determinism: Falseseed: 42data_seed: Noneuse_cpu: Falseaccelerator_config: {'split_batches': False, 'dispatch_batches': None, 'even_batches': True, 'use_seedable_sampler': True, 'non_blocking': False, 'gradient_accumulation_kwargs': None}parallelism_config: Nonedataloader_drop_last: Falsedataloader_num_workers: 0dataloader_pin_memory: Truedataloader_persistent_workers: Falsedataloader_prefetch_factor: Noneremove_unused_columns: Truelabel_names: Nonetrain_sampling_strategy: randomlength_column_name: lengthddp_find_unused_parameters: Noneddp_bucket_cap_mb: Noneddp_broadcast_buffers: Falseddp_static_graph: Noneddp_backend: Noneddp_timeout: 1800fsdp: Nonefsdp_config: Nonedeepspeed: Nonedebug: []skip_memory_metrics: Truedo_predict: Falseresume_from_checkpoint: Nonewarmup_ratio: Nonelocal_rank: -1prompts: Nonebatch_sampler: no_duplicatesmulti_dataset_batch_sampler: proportionalrouter_mapping: {}learning_rate_mapping: {}
Training Logs
| Epoch | Step | Training Loss | comp_nn_accuracy | val_cosine_accuracy |
|---|---|---|---|---|
| -1 | -1 | - | 0.6585 | 0.8995 |
| 0.6173 | 50 | 0.6714 | - | - |
| 1.0 | 81 | - | 0.8171 | 0.9416 |
- The bold row denotes the saved checkpoint.
Training Time
- Training: 1.9 minutes
- Evaluation: 0.8 seconds
- Total: 1.9 minutes
Framework Versions
- Python: 3.12.10
- Sentence Transformers: 5.7.0
- Transformers: 5.14.1
- PyTorch: 2.11.0+cu128
- Accelerate: 1.14.0
- Datasets: 5.0.1
- Tokenizers: 0.22.2
Additional Resources
- Training and Finetuning Embedding Models with Sentence Transformers: the end-to-end guide for training or finetuning Sentence Transformer models.
- Introduction to Matryoshka Embedding Models: variable-size embeddings that can be truncated with minimal quality loss.
- Binary and Scalar Embedding Quantization for Significantly Faster & Cheaper Retrieval: post-training compression of embedding vectors.
- Multimodal Embedding & Reranker Models with Sentence Transformers: use text, image, audio, and video models through the same API.
- Training and Finetuning Multimodal Embedding & Reranker Models with Sentence Transformers: train multimodal embedding models, with a Visual Document Retrieval walkthrough.
Citation
BibTeX
Sentence Transformers
@inproceedings{reimers-2019-sentence-bert,
title = "Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks",
author = "Reimers, Nils and Gurevych, Iryna",
booktitle = "Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing",
month = "11",
year = "2019",
publisher = "Association for Computational Linguistics",
url = "https://arxiv.org/abs/1908.10084",
}
- Downloads last month
- -
Model tree for Aditya109/us-openfda-drug-embed-v1
Base model
FremyCompany/BioLORD-2023Paper for Aditya109/us-openfda-drug-embed-v1
Evaluation results
- Same on compself-reported0.916
- Diff on compself-reported0.514
- Margin on compself-reported0.402
- Margin Norm on compself-reported2.819
- Nn Accuracy on compself-reported0.817
- Cosine Accuracy on valself-reported0.942