SentenceTransformer based on sentence-transformers/LaBSE

This is a sentence-transformers model finetuned from sentence-transformers/LaBSE. It maps sentences & paragraphs to a 768-dimensional dense vector space and can be used for retrieval.

Model Details

Model Description

  • Model Type: Sentence Transformer
  • Base model: sentence-transformers/LaBSE
  • Maximum Sequence Length: 64 tokens
  • Output Dimensionality: 768 dimensions
  • Similarity Function: Cosine Similarity
  • Supported Modality: Text

Model Sources

Full Model Architecture

SentenceTransformer(
  (0): Transformer({'transformer_task': 'feature-extraction', 'modality_config': {'text': {'method': 'forward', 'method_output_name': 'last_hidden_state'}}, 'module_output_name': 'token_embeddings', 'architecture': 'BertModel'})
  (1): Pooling({'embedding_dimension': 768, 'pooling_mode': 'cls', 'include_prompt': True})
  (2): Dense({'in_features': 768, 'out_features': 768, 'bias': True, 'activation_function': 'torch.nn.modules.activation.Tanh', 'module_input_name': 'sentence_embedding', 'module_output_name': 'sentence_embedding'})
  (3): Normalize({})
)

Usage

Direct Usage (Sentence Transformers)

First install the Sentence Transformers library:

pip install -U sentence-transformers

Then you can load this model and run inference.

from sentence_transformers import SentenceTransformer

# Download from the 🤗 Hub
model = SentenceTransformer("sentence_transformers_model_id")
# Run inference
sentences = [
    'ආසන්න වශයෙන් සන්නද්ධ සේවා සහ පොලිස් නිලධාරීන්  75ක් වඩා හොඳ පරිපාලනය සඳහා අනුයුක්ත කරන ලදී.',
    'සිවිල් ආරක්ෂක දෙපාර්තමේන්තුවේ වත්මන් තුන්වන අධ්\u200dයක්ෂ ජනරාල්\xa0චන්ද්\u200dරරත්න පල්ලේගම ( MA, BSc (Hons), PgD, JP (All-Island), FCPM, MAAT (SL) ),\xa0මහතා ශ්\u200dරි ලංකා පරිපාලන සේවයේ (SLAS) විශේෂ ශ්\u200dරේණියේ නිලධාරියෙකි',
    '39,800 කට අධික පිරිසක් සියලු දිස්ත්\u200dරික්ක සහ පළාත්වල සේවය කරන නමුත් , ඉන් බොහෝ පිරිසක් උතුරු හා නැගෙනහිර පළාත්වල, එල්ටීටීඊ ප්\u200dරහාර එල්ල වීමෙන්\xa0පීඩාවට පත් ගම්මානවල සේවයේ යොදවා ඇත .',
]
embeddings = model.encode(sentences)
print(embeddings.shape)
# [3, 768]

# Get the similarity scores for the embeddings
similarities = model.similarity(embeddings, embeddings)
print(similarities)
# tensor([[1.0000, 0.6430, 0.5702],
#         [0.6430, 1.0000, 0.2757],
#         [0.5702, 0.2757, 1.0000]])

Training Details

Training Dataset

Unnamed Dataset

  • Size: 262,664 training samples
  • Columns: sentence_0, sentence_1, and sentence_2
  • Approximate statistics based on the first 100 samples:
    sentence_0 sentence_1 sentence_2
    type string string string
    modality text text text
    details
    • min: 6 tokens
    • mean: 30.52 tokens
    • max: 64 tokens
    • min: 9 tokens
    • mean: 30.84 tokens
    • max: 64 tokens
    • min: 5 tokens
    • mean: 34.91 tokens
    • max: 64 tokens
  • Samples:
    sentence_0 sentence_1 sentence_2
    කෙසේනමුත්, මෙම පංච රථ ඉන්දියානු දේවාල ගෘහනිර්මාණ ශිල්පයේ ප්‍රගමනයට පූර්වාදර්ශයක් වී ඇත. සෙස පංච රථ සතර මෙන් මෙම පාෂාණමය රථය ද මින් පෙර පැවති දැවමය නිර්මාණයක අනුරුවක් විය හැක. සියලුම පංච රථ උතුරු-දකුණු දිශානතිය ඔස්සේ පිහිටා ඇති අතර, පොදු පාදමක පිහිටා ඇත.මේවාට පෙර කිසිදු මේ ආකාරයේ ගෘහනිර්මාණ ක්‍රමවේදයක් දක්නට නොලැබෙන අතර, ඒවා පසුකාලීන විශාල දකුණු ඉන්දියානු ද්‍රවිඩියානු දේවාල ගෘහනිර්මාණ සඳහා "මූලාදර්ශ" වන්නට ඇතැයි විශ්වාස කෙරේ.
    සතර දේවාලයේ කප් සිටුවයි. දිය කපන දිනයේ උගුල්ලා ගඟට විසි කරන්නේ මෙම කපයි.පාන්දර හතරට පමණ කප් සිටුවන අතර අලුත් නුවර සිට කප ගෙන ඒමද සිරිතකි. හය වන දවසේ ඇරඹෙන්නේ කුඹල් පෙරහැරයි.
    විශාල වළාකුලක් එයට සමාන බූ සීමා විශාලත්වයකින් යුතු ඉතා නොගැඹුරු වතුර වලක ඇති තරම් ජලය ඇත. මෙම විද්‍යාවේ වෛද්‍ය විද්‍යාත්මක අතින් වැදගත් වන්නේ වාතය හරහා බෝවන රෝග පිළිබඳ අධ්‍යයනයයි. ඉන් පසුව ජූලි 27 දින ප්‍රාණ රහිත ඔහුගේ දේහය ඇඳ අසල තිබෙනු උපස්ථායකයා විසින් දක්නා ලදී.
  • Loss: MultipleNegativesRankingLoss with these parameters:
    {
        "scale": 20.0,
        "similarity_fct": "cos_sim",
        "gather_across_devices": false,
        "directions": [
            "query_to_doc"
        ],
        "partition_mode": "joint",
        "hardness_mode": null,
        "hardness_strength": 0.0
    }
    

Training Hyperparameters

Non-Default Hyperparameters

  • per_device_train_batch_size: 16
  • num_train_epochs: 1
  • fp16: True
  • per_device_eval_batch_size: 16
  • multi_dataset_batch_sampler: round_robin

All Hyperparameters

Click to expand
  • per_device_train_batch_size: 16
  • num_train_epochs: 1
  • max_steps: -1
  • learning_rate: 5e-05
  • lr_scheduler_type: linear
  • lr_scheduler_kwargs: None
  • warmup_steps: 0
  • optim: adamw_torch
  • optim_args: None
  • weight_decay: 0.0
  • adam_beta1: 0.9
  • adam_beta2: 0.999
  • adam_epsilon: 1e-08
  • optim_target_modules: None
  • gradient_accumulation_steps: 1
  • average_tokens_across_devices: True
  • max_grad_norm: 1
  • label_smoothing_factor: 0.0
  • bf16: False
  • fp16: True
  • bf16_full_eval: False
  • fp16_full_eval: False
  • tf32: None
  • gradient_checkpointing: False
  • gradient_checkpointing_kwargs: None
  • torch_compile: False
  • torch_compile_backend: None
  • torch_compile_mode: None
  • use_liger_kernel: False
  • liger_kernel_config: None
  • use_cache: False
  • neftune_noise_alpha: None
  • torch_empty_cache_steps: None
  • auto_find_batch_size: False
  • log_on_each_node: True
  • logging_nan_inf_filter: True
  • include_num_input_tokens_seen: no
  • log_level: passive
  • log_level_replica: warning
  • disable_tqdm: False
  • project: huggingface
  • trackio_space_id: None
  • trackio_bucket_id: None
  • trackio_static_space_id: None
  • per_device_eval_batch_size: 16
  • prediction_loss_only: True
  • eval_on_start: False
  • eval_do_concat_batches: True
  • eval_use_gather_object: False
  • eval_accumulation_steps: None
  • include_for_metrics: []
  • batch_eval_metrics: False
  • save_only_model: False
  • save_on_each_node: False
  • enable_jit_checkpoint: False
  • push_to_hub: False
  • hub_private_repo: None
  • hub_model_id: None
  • hub_strategy: every_save
  • hub_always_push: False
  • hub_revision: None
  • load_best_model_at_end: False
  • ignore_data_skip: False
  • restore_callback_states_from_checkpoint: False
  • full_determinism: False
  • seed: 42
  • data_seed: None
  • use_cpu: False
  • accelerator_config: {'split_batches': False, 'dispatch_batches': None, 'even_batches': True, 'use_seedable_sampler': True, 'non_blocking': False, 'gradient_accumulation_kwargs': None}
  • parallelism_config: None
  • dataloader_drop_last: False
  • dataloader_num_workers: 0
  • dataloader_pin_memory: True
  • dataloader_persistent_workers: False
  • dataloader_prefetch_factor: None
  • remove_unused_columns: True
  • label_names: None
  • train_sampling_strategy: random
  • length_column_name: length
  • ddp_find_unused_parameters: None
  • ddp_bucket_cap_mb: None
  • ddp_broadcast_buffers: False
  • ddp_static_graph: None
  • ddp_backend: None
  • ddp_timeout: 1800
  • fsdp: None
  • fsdp_config: None
  • deepspeed: None
  • debug: []
  • skip_memory_metrics: True
  • do_predict: False
  • resume_from_checkpoint: None
  • warmup_ratio: None
  • local_rank: -1
  • prompts: None
  • batch_sampler: batch_sampler
  • multi_dataset_batch_sampler: round_robin
  • router_mapping: {}
  • learning_rate_mapping: {}

Training Logs

Click to expand
Epoch Step Training Loss
0.0305 500 1.5952
0.0609 1000 1.1837
0.0914 1500 1.0741
0.1218 2000 1.0447
0.1523 2500 1.0214
0.1827 3000 0.9567
0.2132 3500 0.9446
0.2436 4000 0.9445
0.2741 4500 0.9084
0.3046 5000 0.9049
0.3350 5500 0.8879
0.3655 6000 0.8746
0.3959 6500 0.8342
0.4264 7000 0.8318
0.4568 7500 0.8514
0.4873 8000 0.8166
0.5178 8500 0.8046
0.5482 9000 0.8257
0.5787 9500 0.8034
0.6091 10000 0.7874
0.6396 10500 0.7786
0.6700 11000 0.7563
0.7005 11500 0.7778
0.7309 12000 0.7434
0.7614 12500 0.7495
0.7919 13000 0.7639
0.8223 13500 0.7398
0.8528 14000 0.7493
0.8832 14500 0.7284
0.9137 15000 0.7317
0.9441 15500 0.7351
0.9746 16000 0.7062
0.0305 500 0.6713
0.0609 1000 0.6940
0.0914 1500 0.6673
0.1218 2000 0.6850
0.1523 2500 0.7092
0.1827 3000 0.6829
0.2132 3500 0.6903
0.2436 4000 0.6828
0.2741 4500 0.6641
0.3046 5000 0.6574
0.3350 5500 0.6672
0.3655 6000 0.6652
0.3959 6500 0.6573
0.4264 7000 0.6644
0.4568 7500 0.6387
0.4873 8000 0.6370
0.5178 8500 0.6411
0.5482 9000 0.6358
0.5787 9500 0.6207
0.6091 10000 0.6037
0.6396 10500 0.6328
0.6700 11000 0.6017
0.7005 11500 0.6248
0.7309 12000 0.5891
0.7614 12500 0.5956
0.7919 13000 0.5924
0.8223 13500 0.5777
0.8528 14000 0.5891
0.8832 14500 0.5714
0.9137 15000 0.5815
0.9441 15500 0.5812
0.9746 16000 0.5578
0.0305 500 0.1675
0.0609 1000 0.1821
0.0914 1500 0.1734
0.1218 2000 0.1907
0.1523 2500 0.2056
0.1827 3000 0.1925
0.2132 3500 0.2149
0.2436 4000 0.2120
0.2741 4500 0.2132
0.3046 5000 0.2159
0.3350 5500 0.2291
0.3655 6000 0.2418
0.3959 6500 0.2509
0.4264 7000 0.2597
0.4568 7500 0.2718
0.4873 8000 0.2700
0.5178 8500 0.2905
0.5482 9000 0.3004
0.5787 9500 0.3076
0.6091 10000 0.3025
0.6396 10500 0.3435
0.6700 11000 0.3351
0.7005 11500 0.3889
0.7309 12000 0.3750
0.7614 12500 0.3977
0.7919 13000 0.4149
0.8223 13500 0.4214
0.8528 14000 0.4504
0.8832 14500 0.4687
0.9137 15000 0.4966
0.9441 15500 0.5326
0.9746 16000 0.5350
0.0305 500 0.0357
0.0609 1000 0.0421
0.0914 1500 0.0443
0.1218 2000 0.0545
0.1523 2500 0.0527
0.1827 3000 0.0509
0.2132 3500 0.0579
0.2436 4000 0.0543
0.2741 4500 0.0620
0.3046 5000 0.0640
0.3350 5500 0.0656
0.3655 6000 0.0689
0.3959 6500 0.0721
0.4264 7000 0.0798
0.4568 7500 0.0835
0.4873 8000 0.0870
0.5178 8500 0.0944
0.5482 9000 0.1082
0.5787 9500 0.1137
0.6091 10000 0.1115
0.6396 10500 0.1397
0.6700 11000 0.1484
0.7005 11500 0.1895
0.7309 12000 0.1975
0.7614 12500 0.2279
0.7919 13000 0.2547
0.8223 13500 0.2777
0.8528 14000 0.3216
0.8832 14500 0.3612
0.9137 15000 0.4126
0.9441 15500 0.4811
0.9746 16000 0.5160
0.0305 500 0.0112
0.0609 1000 0.0173
0.0914 1500 0.0166
0.1218 2000 0.0204
0.1523 2500 0.0221
0.1827 3000 0.0197
0.2132 3500 0.0231
0.2436 4000 0.0218
0.2741 4500 0.0234
0.3046 5000 0.0245
0.3350 5500 0.0237
0.3655 6000 0.0255
0.3959 6500 0.0262
0.4264 7000 0.0298
0.4568 7500 0.0337
0.4873 8000 0.0332
0.5178 8500 0.0334
0.5482 9000 0.0406
0.5787 9500 0.0451
0.6091 10000 0.0445
0.6396 10500 0.0568
0.6700 11000 0.0589
0.7005 11500 0.0869
0.7309 12000 0.0951
0.7614 12500 0.1140
0.7919 13000 0.1461
0.8223 13500 0.1731
0.8528 14000 0.2140
0.8832 14500 0.2686
0.9137 15000 0.3328
0.9441 15500 0.4245
0.9746 16000 0.5006
0.0305 500 0.0048
0.0609 1000 0.0107
0.0914 1500 0.0111
0.1218 2000 0.0112
0.1523 2500 0.0104
0.1827 3000 0.0111
0.2132 3500 0.0114
0.2436 4000 0.0108
0.2741 4500 0.0107
0.3046 5000 0.0127
0.3350 5500 0.0132
0.3655 6000 0.0132
0.3959 6500 0.0140
0.4264 7000 0.0138
0.4568 7500 0.0164
0.4873 8000 0.0173
0.5178 8500 0.0159
0.5482 9000 0.0197
0.5787 9500 0.0205
0.6091 10000 0.0212
0.6396 10500 0.0261
0.6700 11000 0.0274
0.7005 11500 0.0413
0.7309 12000 0.0509
0.7614 12500 0.0586
0.7919 13000 0.0796
0.8223 13500 0.1056
0.8528 14000 0.1412
0.8832 14500 0.1975
0.9137 15000 0.2672
0.9441 15500 0.3724
0.9746 16000 0.4738
0.0305 500 0.0024
0.0609 1000 0.0040
0.0914 1500 0.0034
0.1218 2000 0.0060
0.1523 2500 0.0082
0.1827 3000 0.0075
0.2132 3500 0.0074
0.2436 4000 0.0069
0.2741 4500 0.0082
0.3046 5000 0.0076
0.3350 5500 0.0100
0.3655 6000 0.0084
0.3959 6500 0.0093
0.4264 7000 0.0097
0.4568 7500 0.0104
0.4873 8000 0.0097
0.5178 8500 0.0100
0.5482 9000 0.0119
0.5787 9500 0.0122
0.6091 10000 0.0138
0.6396 10500 0.0155
0.6700 11000 0.0186
0.7005 11500 0.0242
0.7309 12000 0.0276
0.7614 12500 0.0359
0.7919 13000 0.0489
0.8223 13500 0.0683
0.8528 14000 0.0981
0.8832 14500 0.1450
0.9137 15000 0.2173
0.9441 15500 0.3247
0.9746 16000 0.4529
0.0305 500 0.0025
0.0609 1000 0.0042
0.0914 1500 0.0038
0.1218 2000 0.0057
0.1523 2500 0.0068
0.1827 3000 0.0052
0.2132 3500 0.0058
0.2436 4000 0.0056
0.2741 4500 0.0058
0.3046 5000 0.0047
0.3350 5500 0.0056
0.3655 6000 0.0065
0.3959 6500 0.0069
0.4264 7000 0.0054
0.4568 7500 0.0060
0.4873 8000 0.0065
0.5178 8500 0.0063
0.5482 9000 0.0077
0.5787 9500 0.0081
0.6091 10000 0.0077
0.6396 10500 0.0098
0.6700 11000 0.0116
0.7005 11500 0.0149
0.7309 12000 0.0177
0.7614 12500 0.0188
0.7919 13000 0.0301
0.8223 13500 0.0415
0.8528 14000 0.0631
0.8832 14500 0.1010
0.9137 15000 0.1621
0.9441 15500 0.2720
0.9746 16000 0.4265

Training Time

  • Training: 56.5 minutes

Evaluation

Evaluated on a held-out split of topically-coherent sentence pairs (positives) against paragraph-boundary hard negatives, used as the coherence signal in akshara-kit's neuro-symbolic chunker:

AUC mean sim, coherent mean sim, hard-negative
Base LaBSE 0.7269 0.3871 0.2746
Fine-tuned (this model) 0.8832 0.5313 0.1620

Bootstrap 95% CI on the AUC improvement over base LaBSE: [+0.1419, +0.1694].

Framework Versions

  • Python: 3.10.12
  • Sentence Transformers: 5.6.1
  • Transformers: 5.14.1
  • PyTorch: 2.5.1+cu121
  • Accelerate: 1.14.0
  • Datasets: 5.0.1
  • Tokenizers: 0.22.2

Citation

BibTeX

Sentence Transformers

@inproceedings{reimers-2019-sentence-bert,
    title = "Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks",
    author = "Reimers, Nils and Gurevych, Iryna",
    booktitle = "Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing",
    month = "11",
    year = "2019",
    publisher = "Association for Computational Linguistics",
    url = "https://arxiv.org/abs/1908.10084",
}

MultipleNegativesRankingLoss

@misc{oord2019representationlearningcontrastivepredictive,
      title={Representation Learning with Contrastive Predictive Coding},
      author={Aaron van den Oord and Yazhe Li and Oriol Vinyals},
      year={2019},
      eprint={1807.03748},
      archivePrefix={arXiv},
      primaryClass={cs.LG},
      url={https://arxiv.org/abs/1807.03748},
}
Downloads last month
80
Safetensors
Model size
0.5B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Nimsara2001/labse-sinhala-finetuned

Finetuned
(94)
this model

Papers for Nimsara2001/labse-sinhala-finetuned