SentenceTransformer based on sentence-transformers/all-mpnet-base-v2

This is a sentence-transformers model finetuned from sentence-transformers/all-mpnet-base-v2. It maps inputs to a 768-dimensional dense vector space and can be used for semantic textual similarity, semantic search, paraphrase mining, classification, clustering, and more.

Model Details

Model Description

  • Model Type: Sentence Transformer
  • Base model: sentence-transformers/all-mpnet-base-v2
  • Maximum Sequence Length: 32 tokens
  • Output Dimensionality: 768 dimensions
  • Similarity Function: Cosine Similarity
  • Supported Modality: Text

Model Sources

Full Model Architecture

SentenceTransformer(
  (0): Transformer({'transformer_task': 'feature-extraction', 'modality_config': {'text': {'method': 'forward', 'method_output_name': 'last_hidden_state'}}, 'module_output_name': 'token_embeddings', 'architecture': 'MPNetModel'})
  (1): Pooling({'embedding_dimension': 768, 'pooling_mode': 'mean', 'include_prompt': True})
  (2): Normalize({'module_input_name': 'sentence_embedding', 'module_output_name': 'sentence_embedding'})
)

Usage

Direct Usage (Sentence Transformers)

First install the Sentence Transformers library:

pip install -U sentence-transformers

Then you can load this model and run inference.

from sentence_transformers import SentenceTransformer

# Download from the 🤗 Hub
model = SentenceTransformer("seanfarrell/pettag_emed_train_test")
# Run inference
sentences = [
    'Pythiosis',
    'pythiosis',
    'Nasal step',
]
embeddings = model.encode(sentences)
print(embeddings.shape)
# [3, 768]

# Get the similarity scores for the embeddings
similarities = model.similarity(embeddings, embeddings)
print(similarities)
# tensor([[1.0000, 1.0000, 0.0459],
#         [1.0000, 1.0000, 0.0459],
#         [0.0459, 0.0459, 1.0000]])

Training Details

Training Dataset

Unnamed Dataset

  • Size: 10,000,000 training samples
  • Columns: sentence_0 and sentence_1
  • Approximate statistics based on the first 100 samples:
    sentence_0 sentence_1
    type string string
    modality text text
    details
    • min: 5 tokens
    • mean: 11.07 tokens
    • max: 27 tokens
    • min: 5 tokens
    • mean: 11.6 tokens
    • max: 27 tokens
  • Samples:
    sentence_0 sentence_1
    other specified structural developmental anomalies of cervix uteri Other specified structural developmental anomalies of cervix uteri
    baroreflex failure Baroreflex failure
    Other specified portal hypertension banti syndrome
  • Loss: MultipleNegativesRankingLoss with these parameters:
    {
        "scale": 20.0,
        "similarity_fct": "cos_sim",
        "gather_across_devices": false,
        "directions": [
            "query_to_doc"
        ],
        "partition_mode": "joint",
        "hardness_mode": null,
        "hardness_strength": 0.0
    }
    

Training Hyperparameters

Non-Default Hyperparameters

  • per_device_train_batch_size: 512
  • fp16: True
  • per_device_eval_batch_size: 512
  • multi_dataset_batch_sampler: round_robin

All Hyperparameters

Click to expand
  • per_device_train_batch_size: 512
  • num_train_epochs: 3
  • max_steps: -1
  • learning_rate: 5e-05
  • lr_scheduler_type: linear
  • lr_scheduler_kwargs: None
  • warmup_steps: 0
  • optim: adamw_torch_fused
  • optim_args: None
  • weight_decay: 0.0
  • adam_beta1: 0.9
  • adam_beta2: 0.999
  • adam_epsilon: 1e-08
  • optim_target_modules: None
  • gradient_accumulation_steps: 1
  • average_tokens_across_devices: True
  • max_grad_norm: 1
  • label_smoothing_factor: 0.0
  • bf16: False
  • fp16: True
  • bf16_full_eval: False
  • fp16_full_eval: False
  • tf32: None
  • gradient_checkpointing: False
  • gradient_checkpointing_kwargs: None
  • torch_compile: False
  • torch_compile_backend: None
  • torch_compile_mode: None
  • use_liger_kernel: False
  • liger_kernel_config: None
  • use_cache: False
  • neftune_noise_alpha: None
  • torch_empty_cache_steps: None
  • auto_find_batch_size: False
  • log_on_each_node: True
  • logging_nan_inf_filter: True
  • include_num_input_tokens_seen: no
  • log_level: passive
  • log_level_replica: warning
  • disable_tqdm: False
  • project: huggingface
  • trackio_space_id: None
  • trackio_bucket_id: None
  • trackio_static_space_id: None
  • per_device_eval_batch_size: 512
  • prediction_loss_only: True
  • eval_on_start: False
  • eval_do_concat_batches: True
  • eval_use_gather_object: False
  • eval_accumulation_steps: None
  • include_for_metrics: []
  • batch_eval_metrics: False
  • save_only_model: False
  • save_on_each_node: False
  • enable_jit_checkpoint: False
  • push_to_hub: False
  • hub_private_repo: None
  • hub_model_id: None
  • hub_strategy: every_save
  • hub_always_push: False
  • hub_revision: None
  • load_best_model_at_end: False
  • ignore_data_skip: False
  • restore_callback_states_from_checkpoint: False
  • full_determinism: False
  • seed: 42
  • data_seed: None
  • use_cpu: False
  • accelerator_config: {'split_batches': False, 'dispatch_batches': None, 'even_batches': True, 'use_seedable_sampler': True, 'non_blocking': False, 'gradient_accumulation_kwargs': None}
  • parallelism_config: None
  • dataloader_drop_last: False
  • dataloader_num_workers: 0
  • dataloader_pin_memory: True
  • dataloader_persistent_workers: False
  • dataloader_prefetch_factor: None
  • dataloader_multiprocessing_context: None
  • dataloader_in_order: True
  • remove_unused_columns: True
  • label_names: None
  • train_sampling_strategy: random
  • length_column_name: length
  • ddp_find_unused_parameters: None
  • ddp_bucket_cap_mb: None
  • ddp_broadcast_buffers: False
  • ddp_static_graph: None
  • ddp_backend: None
  • ddp_timeout: 1800
  • fsdp: None
  • fsdp_config: None
  • deepspeed: None
  • debug: []
  • skip_memory_metrics: True
  • do_predict: False
  • resume_from_checkpoint: None
  • local_rank: -1
  • prompts: None
  • batch_sampler: batch_sampler
  • multi_dataset_batch_sampler: round_robin
  • router_mapping: {}
  • learning_rate_mapping: {}
  • warmup_ratio: None

Training Logs

Click to expand
Epoch Step Training Loss
0.0256 500 0.8003
0.0512 1000 0.5966
0.0768 1500 0.4501
0.1024 2000 0.3559
0.1280 2500 0.2971
0.1536 3000 0.2566
0.1792 3500 0.2210
0.2048 4000 0.1964
0.2304 4500 0.1773
0.2560 5000 0.1612
0.2816 5500 0.1485
0.3072 6000 0.1390
0.3328 6500 0.1301
0.3584 7000 0.1200
0.3840 7500 0.1129
0.4096 8000 0.1095
0.4352 8500 0.1031
0.4608 9000 0.0983
0.4864 9500 0.0930
0.5120 10000 0.0883
0.5376 10500 0.0865
0.5632 11000 0.0854
0.5888 11500 0.0815
0.6144 12000 0.0783
0.6400 12500 0.0764
0.6656 13000 0.0737
0.6912 13500 0.0745
0.7168 14000 0.0713
0.7424 14500 0.0705
0.7680 15000 0.0682
0.7936 15500 0.0674
0.8192 16000 0.0650
0.8448 16500 0.0646
0.8704 17000 0.0640
0.8960 17500 0.0624
0.9216 18000 0.0604
0.9472 18500 0.0610
0.9728 19000 0.0596
0.9984 19500 0.0599
1.0240 20000 0.0562
1.0496 20500 0.0570
1.0752 21000 0.0563
1.1008 21500 0.0550
1.1264 22000 0.0565
1.1520 22500 0.0551
1.1776 23000 0.0549
1.2032 23500 0.0538
1.2288 24000 0.0532
1.2544 24500 0.0527
1.2800 25000 0.0538
1.3055 25500 0.0514
1.3311 26000 0.0521
1.3567 26500 0.0511
1.3823 27000 0.0504
1.4079 27500 0.0501
1.4335 28000 0.0503
1.4591 28500 0.0496
1.4847 29000 0.0509
1.5103 29500 0.0492
1.5359 30000 0.0488
1.5615 30500 0.0489
1.5871 31000 0.0473
1.6127 31500 0.0483
1.6383 32000 0.0477
1.6639 32500 0.0478
1.6895 33000 0.0473
1.7151 33500 0.0471
1.7407 34000 0.0469
1.7663 34500 0.0467
1.7919 35000 0.0449
1.8175 35500 0.0471
1.8431 36000 0.0462
1.8687 36500 0.0463
1.8943 37000 0.0457
1.9199 37500 0.0458
1.9455 38000 0.0465
1.9711 38500 0.0458
1.9967 39000 0.0453
2.0223 39500 0.0443
2.0479 40000 0.0444
2.0735 40500 0.0444
2.0991 41000 0.0450
2.1247 41500 0.0446
2.1503 42000 0.0431
2.1759 42500 0.0437
2.2015 43000 0.0446
2.2271 43500 0.0440
2.2527 44000 0.0430
2.2783 44500 0.0440
2.3039 45000 0.0446
2.3295 45500 0.0441
2.3551 46000 0.0423
2.3807 46500 0.0428
2.4063 47000 0.0429
2.4319 47500 0.0425
2.4575 48000 0.0422
2.4831 48500 0.0426
2.5087 49000 0.0421
2.5343 49500 0.0422
2.5599 50000 0.0423
2.5855 50500 0.0426
2.6111 51000 0.0413
2.6367 51500 0.0413
2.6623 52000 0.0413
2.6879 52500 0.0419
2.7135 53000 0.0407
2.7391 53500 0.0425
2.7647 54000 0.0409
2.7903 54500 0.0415
2.8159 55000 0.0404
2.8415 55500 0.0418
2.8671 56000 0.0417
2.8927 56500 0.0408
2.9183 57000 0.0412
2.9439 57500 0.0405
2.9695 58000 0.0412
2.9951 58500 0.0412

Training Time

  • Training: 4.3 hours

Framework Versions

  • Python: 3.12.3
  • Sentence Transformers: 6.1.0
  • Transformers: 5.19.0
  • PyTorch: 2.14.1+cu130
  • Accelerate: 1.15.0
  • Datasets: 5.1.0
  • Tokenizers: 0.23.3

Additional Resources

Citation

BibTeX

Sentence Transformers

@inproceedings{reimers-2019-sentence-bert,
    title = "Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks",
    author = "Reimers, Nils and Gurevych, Iryna",
    booktitle = "Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing",
    month = "11",
    year = "2019",
    publisher = "Association for Computational Linguistics",
    url = "https://arxiv.org/abs/1908.10084",
}

MultipleNegativesRankingLoss

@misc{oord2019representationlearningcontrastivepredictive,
      title={Representation Learning with Contrastive Predictive Coding},
      author={Aaron van den Oord and Yazhe Li and Oriol Vinyals},
      year={2019},
      eprint={1807.03748},
      archivePrefix={arXiv},
      primaryClass={cs.LG},
      url={https://arxiv.org/abs/1807.03748},
}
Downloads last month
12
Safetensors
Model size
0.1B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for seanfarrell/pettag_emed_train_test

Finetuned
(401)
this model

Papers for seanfarrell/pettag_emed_train_test