SentenceTransformer based on sentence-transformers/paraphrase-multilingual-MiniLM-L12-v2

This is a sentence-transformers model finetuned from sentence-transformers/paraphrase-multilingual-MiniLM-L12-v2. It maps sentences & paragraphs to a 384-dimensional dense vector space and can be used for retrieval.

Model Details

Model Description

Model Sources

Full Model Architecture

SentenceTransformer(
  (0): Transformer({'transformer_task': 'feature-extraction', 'modality_config': {'text': {'method': 'forward', 'method_output_name': 'last_hidden_state'}}, 'module_output_name': 'token_embeddings', 'architecture': 'BertModel'})
  (1): Pooling({'embedding_dimension': 384, 'pooling_mode': 'mean', 'include_prompt': True})
)

Usage

Direct Usage (Sentence Transformers)

First install the Sentence Transformers library:

pip install -U sentence-transformers

Then you can load this model and run inference.

from sentence_transformers import SentenceTransformer

# Download from the 🤗 Hub
model = SentenceTransformer("sentence_transformers_model_id")
# Run inference
sentences = [
    'anāpatti hi so rukkho, hoti ekakulassa ce.',
    'there is indeed no offense if that tree belongs to a single family.',
    'there are four kinds of purity: purity of instruction, purity of restraint, purity of seeking, and purity of reflection.',
]
embeddings = model.encode(sentences)
print(embeddings.shape)
# [3, 384]

# Get the similarity scores for the embeddings
similarities = model.similarity(embeddings, embeddings)
print(similarities)
# tensor([[ 1.0000,  0.8338, -0.1214],
#         [ 0.8338,  1.0000, -0.1454],
#         [-0.1214, -0.1454,  1.0000]])

Training Details

Training Dataset

Unnamed Dataset

  • Size: 844,384 training samples
  • Columns: sentence_0 and sentence_1
  • Approximate statistics based on the first 100 samples:
    sentence_0 sentence_1
    type string string
    modality text text
    details
    • min: 11 tokens
    • mean: 40.61 tokens
    • max: 128 tokens
    • min: 11 tokens
    • mean: 38.81 tokens
    • max: 128 tokens
  • Samples:
    sentence_0 sentence_1
    tasi alaṅkāre, bhūvādi. tasi [is used] in the sense of adorning; it belongs to the bhūvādi group.
    evaṃ phussena yutto māso phusso, maghāya yutto māso māgho, phagguniyā yutto māso phagguno, cittāya yutto māso citto, visākhāya yutto māso vesākho, jeṭṭhāya yutto māso jeṭṭho, uttarāsāḷhāya yutto māso āsāḷho, āsāḷhī vā, savaṇena yutto māso sāvaṇo, sāvaṇī. similarly, a month conjoined with phussa (pusya) is phusso; a month conjoined with maghā is māgho; a month conjoined with phaggunī (phalgunī) is phagguno; a month conjoined with cittā is citto; a month conjoined with visākhā is vesākho; a month conjoined with jeṭṭhā (jyesthā) is jeṭṭho; a month conjoined with uttarāsāḷhā (uttarāṣāḍhā) is āsāḷho or āsāḷhī; a month conjoined with savaṇa (śravaṇa) is sāvaṇo or sāvaṇī.
    īādimhi-akari, kari, saṅkhari, abhisaṅkhari, akubbi, kubbi, akrubbi, krubbi, akayiri, kayiri, akaruṃ, karuṃ, saṅkharuṃ, abhi, saṅkharuṃ, akariṃsu, kariṃsu, saṅkhariṃsu, abhisaṅkhariṃsu, akubbiṃsu, kubbiṃsu, akrubbiṃsu, krubbiṃsu, akayiriṃsu, kayiriṃsu, akayiruṃ, kayiruṃ. in the past tense (ī-ādi): akari, kari, saṅkhari, abhisaṅkhari, akubbi, kubbi, akrubbi, krubbi, akayiri, kayiri, akaruṃ, karuṃ, saṅkharuṃ, abhisaṅkharuṃ, akariṃsu, kariṃsu, saṅkhariṃsu, abhisaṅkhariṃsu, akubbiṃsu, kubbiṃsu, akrubbiṃsu, krubbiṃsu, akayiriṃsu, kayiriṃsu, akayiruṃ, kayiruṃ.
  • Loss: MultipleNegativesRankingLoss with these parameters:
    {
        "scale": 20.0,
        "similarity_fct": "cos_sim",
        "gather_across_devices": false,
        "directions": [
            "query_to_doc"
        ],
        "partition_mode": "joint",
        "hardness_mode": null,
        "hardness_strength": 0.0
    }
    

Training Hyperparameters

Non-Default Hyperparameters

  • per_device_train_batch_size: 64
  • num_train_epochs: 1
  • per_device_eval_batch_size: 64
  • multi_dataset_batch_sampler: round_robin

All Hyperparameters

Click to expand
  • per_device_train_batch_size: 64
  • num_train_epochs: 1
  • max_steps: -1
  • learning_rate: 5e-05
  • lr_scheduler_type: linear
  • lr_scheduler_kwargs: None
  • warmup_steps: 0
  • optim: adamw_torch_fused
  • optim_args: None
  • weight_decay: 0.0
  • adam_beta1: 0.9
  • adam_beta2: 0.999
  • adam_epsilon: 1e-08
  • optim_target_modules: None
  • gradient_accumulation_steps: 1
  • average_tokens_across_devices: True
  • max_grad_norm: 1
  • label_smoothing_factor: 0.0
  • bf16: False
  • fp16: False
  • bf16_full_eval: False
  • fp16_full_eval: False
  • tf32: None
  • gradient_checkpointing: False
  • gradient_checkpointing_kwargs: None
  • torch_compile: False
  • torch_compile_backend: None
  • torch_compile_mode: None
  • use_liger_kernel: False
  • liger_kernel_config: None
  • use_cache: False
  • neftune_noise_alpha: None
  • torch_empty_cache_steps: None
  • auto_find_batch_size: False
  • log_on_each_node: True
  • logging_nan_inf_filter: True
  • include_num_input_tokens_seen: no
  • log_level: passive
  • log_level_replica: warning
  • disable_tqdm: False
  • project: huggingface
  • trackio_space_id: None
  • trackio_bucket_id: None
  • trackio_static_space_id: None
  • per_device_eval_batch_size: 64
  • prediction_loss_only: True
  • eval_on_start: False
  • eval_do_concat_batches: True
  • eval_use_gather_object: False
  • eval_accumulation_steps: None
  • include_for_metrics: []
  • batch_eval_metrics: False
  • save_only_model: False
  • save_on_each_node: False
  • enable_jit_checkpoint: False
  • push_to_hub: False
  • hub_private_repo: None
  • hub_model_id: None
  • hub_strategy: every_save
  • hub_always_push: False
  • hub_revision: None
  • load_best_model_at_end: False
  • ignore_data_skip: False
  • restore_callback_states_from_checkpoint: False
  • full_determinism: False
  • seed: 42
  • data_seed: None
  • use_cpu: False
  • accelerator_config: {'split_batches': False, 'dispatch_batches': None, 'even_batches': True, 'use_seedable_sampler': True, 'non_blocking': False, 'gradient_accumulation_kwargs': None}
  • parallelism_config: None
  • dataloader_drop_last: False
  • dataloader_num_workers: 0
  • dataloader_pin_memory: True
  • dataloader_persistent_workers: False
  • dataloader_prefetch_factor: None
  • remove_unused_columns: True
  • label_names: None
  • train_sampling_strategy: random
  • length_column_name: length
  • ddp_find_unused_parameters: None
  • ddp_bucket_cap_mb: None
  • ddp_broadcast_buffers: False
  • ddp_static_graph: None
  • ddp_backend: None
  • ddp_timeout: 1800
  • fsdp: []
  • fsdp_config: {'min_num_params': 0, 'xla': False, 'xla_fsdp_v2': False, 'xla_fsdp_grad_ckpt': False}
  • deepspeed: None
  • debug: []
  • skip_memory_metrics: True
  • do_predict: False
  • resume_from_checkpoint: None
  • warmup_ratio: None
  • local_rank: -1
  • prompts: None
  • batch_sampler: batch_sampler
  • multi_dataset_batch_sampler: round_robin
  • router_mapping: {}
  • learning_rate_mapping: {}

Training Logs

Epoch Step Training Loss
0.0379 500 2.8581
0.0758 1000 0.8063
0.1137 1500 0.4569
0.1516 2000 0.2829
0.1895 2500 0.1918
0.2274 3000 0.1640
0.2653 3500 0.1384
0.3032 4000 0.1248
0.3411 4500 0.1068
0.3790 5000 0.0975
0.4169 5500 0.0925
0.4548 6000 0.0888
0.4926 6500 0.0825
0.5305 7000 0.0800
0.5684 7500 0.0713
0.6063 8000 0.0734
0.6442 8500 0.0729
0.6821 9000 0.0646
0.7200 9500 0.0676
0.7579 10000 0.0633
0.7958 10500 0.0595
0.8337 11000 0.0589
0.8716 11500 0.0567
0.9095 12000 0.0594
0.9474 12500 0.0542
0.9853 13000 0.0559

Training Time

  • Training: 3.1 hours

Framework Versions

  • Python: 3.12.13
  • Sentence Transformers: 5.5.1
  • Transformers: 5.9.0
  • PyTorch: 2.10.0+cu128
  • Accelerate: 1.13.0
  • Datasets: 4.8.5
  • Tokenizers: 0.22.2

Citation

BibTeX

Sentence Transformers

@inproceedings{reimers-2019-sentence-bert,
    title = "Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks",
    author = "Reimers, Nils and Gurevych, Iryna",
    booktitle = "Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing",
    month = "11",
    year = "2019",
    publisher = "Association for Computational Linguistics",
    url = "https://arxiv.org/abs/1908.10084",
}

MultipleNegativesRankingLoss

@misc{oord2019representationlearningcontrastivepredictive,
      title={Representation Learning with Contrastive Predictive Coding},
      author={Aaron van den Oord and Yazhe Li and Oriol Vinyals},
      year={2019},
      eprint={1807.03748},
      archivePrefix={arXiv},
      primaryClass={cs.LG},
      url={https://arxiv.org/abs/1807.03748},
}
Downloads last month
37
Safetensors
Model size
0.1B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for dhammanana/Tipitaka_MiniLM-L12

Papers for dhammanana/Tipitaka_MiniLM-L12