Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks
Paper • 1908.10084 • Published • 16
How to use AniruddhaAI/sanskrit-e5-small-retrieval with sentence-transformers:
from sentence_transformers import SentenceTransformer
model = SentenceTransformer("AniruddhaAI/sanskrit-e5-small-retrieval")
sentences = [
"query: The two princes Vinda and Anuvinda, both fierce bowmen, for the sake of their friend Duryodhana, rushed upon Virata the ruler of the Matsyas, heedless of their lives and holding their weapons over head.",
"passage: Invite them with plenty of seats and welcome them; for they remaining unseated, who is capable of taking his seat?",
"passage: The Sankhya system, the Aranyaka-Veda, and the Pancharatra scriptures, are all identical and form parts of one whole. This is the religion of those who are devoted whole. mindedly to Narayana—the religion that has Narayana for its Soul.",
"passage: विन्दानुविन्दावावन्त्यौविराटं मस्यमार्च्छताम्। प्राणांस्त्यक्त्वा महेष्वासौ मित्रार्थेऽभ्युद्यतायुधौ॥"
]
embeddings = model.encode(sentences)
similarities = model.similarity(embeddings, embeddings)
print(similarities.shape)
# [4, 4]This is a sentence-transformers model finetuned from intfloat/multilingual-e5-small. It maps sentences & paragraphs to a 384-dimensional dense vector space and can be used for retrieval.
SentenceTransformer(
(0): Transformer({'transformer_task': 'feature-extraction', 'modality_config': {'text': {'method': 'forward', 'method_output_name': 'last_hidden_state'}}, 'module_output_name': 'token_embeddings', 'architecture': 'BertModel'})
(1): Pooling({'embedding_dimension': 384, 'pooling_mode': 'mean', 'include_prompt': True})
(2): Normalize({})
)
First install the Sentence Transformers library:
pip install -U sentence-transformers
Then you can load this model and run inference.
from sentence_transformers import SentenceTransformer
# Download from the 🤗 Hub
model = SentenceTransformer("AniruddhaAI/sanskrit-e5-small-retrieval")
# Run inference
sentences = [
'query: He will certainly accomplish that which is pleasing to you. Surrounding by his brave sons all of whom possess arms like maces. O Brahmana, give me leave to depart, for I have now abandoned all weapons."',
'passage: प्रियं च ते सर्वमेतत् करिष्यति न संशयः॥ पुत्रैः परिवृतः सर्वैः शूरैः परिघबाहुभिः। विसर्जयस्व मां ब्रह्मन् न्यस्तशस्त्रोऽस्मि साम्प्रतम्॥',
'passage: Fearing death in seasons of distress, the Vishvedevas, the Saddhyas, the Brahmanas, and great Rishis, do not hesitate to follow the alternative provisions laid down in the scriptures.',
]
embeddings = model.encode(sentences)
print(embeddings.shape)
# [3, 384]
# Get the similarity scores for the embeddings
similarities = model.similarity(embeddings, embeddings)
print(similarities)
# tensor([[ 1.0000, 0.7581, -0.0554],
# [ 0.7581, 1.0000, 0.0291],
# [-0.0554, 0.0291, 1.0000]])
sentence_0 and sentence_1| sentence_0 | sentence_1 | |
|---|---|---|
| type | string | string |
| modality | text | text |
| details |
|
|
| sentence_0 | sentence_1 |
|---|---|
query: Steeds, of the huge of donkeys, with backs of the hue of mice and with necks proudly drawn up, bore Vyaghradatta. |
passage: रासभारुणवर्णाभाः पृष्ठतो मूषिकप्रभाः। वल्गन्त इव संयत्ता व्याघ्रदत्तमुदावहन्॥ |
query: क्रौञ्चं तु गिरिमासाद्य बिलं तस्य सुदुर्गमम् । अप्रमत्तैः प्रवेष्टव्यं दुष्प्रवेशं हि तत् स्मृतम्॥ |
passage: And coming to the Kruncha mountain, you should, having your wits about you, enter its inaccessible cavern; for that is well known as difficult of entrance. |
query: Hearing these words of the vulture, the grief of the kinsmen seemed to decrease, and placing the child on the naked earth they were about to go away. |
passage: ततो गृध्रवचः श्रुत्वा प्राक्रोशन्तस्तदा नृप। बान्धवास्तेऽभ्यगच्छन्त पुत्रमुत्सृज्य भूतले॥ |
MultipleNegativesRankingLoss with these parameters:{
"scale": 20.0,
"similarity_fct": "cos_sim",
"gather_across_devices": false,
"directions": [
"query_to_doc"
],
"partition_mode": "joint",
"hardness_mode": null,
"hardness_strength": 0.0
}
per_device_train_batch_size: 64num_train_epochs: 1fp16: Trueper_device_eval_batch_size: 64multi_dataset_batch_sampler: round_robinper_device_train_batch_size: 64num_train_epochs: 1max_steps: -1learning_rate: 5e-05lr_scheduler_type: linearlr_scheduler_kwargs: Nonewarmup_steps: 0optim: adamw_torch_fusedoptim_args: Noneweight_decay: 0.0adam_beta1: 0.9adam_beta2: 0.999adam_epsilon: 1e-08optim_target_modules: Nonegradient_accumulation_steps: 1average_tokens_across_devices: Truemax_grad_norm: 1label_smoothing_factor: 0.0bf16: Falsefp16: Truebf16_full_eval: Falsefp16_full_eval: Falsetf32: Nonegradient_checkpointing: Falsegradient_checkpointing_kwargs: Nonetorch_compile: Falsetorch_compile_backend: Nonetorch_compile_mode: Noneuse_liger_kernel: Falseliger_kernel_config: Noneuse_cache: Falseneftune_noise_alpha: Nonetorch_empty_cache_steps: Noneauto_find_batch_size: Falselog_on_each_node: Truelogging_nan_inf_filter: Trueinclude_num_input_tokens_seen: nolog_level: passivelog_level_replica: warningdisable_tqdm: Falseproject: huggingfacetrackio_space_id: Nonetrackio_bucket_id: Nonetrackio_static_space_id: Noneper_device_eval_batch_size: 64prediction_loss_only: Trueeval_on_start: Falseeval_do_concat_batches: Trueeval_use_gather_object: Falseeval_accumulation_steps: Noneinclude_for_metrics: []batch_eval_metrics: Falsesave_only_model: Falsesave_on_each_node: Falseenable_jit_checkpoint: Falsepush_to_hub: Falsehub_private_repo: Nonehub_model_id: Nonehub_strategy: every_savehub_always_push: Falsehub_revision: Noneload_best_model_at_end: Falseignore_data_skip: Falserestore_callback_states_from_checkpoint: Falsefull_determinism: Falseseed: 42data_seed: Noneuse_cpu: Falseaccelerator_config: {'split_batches': False, 'dispatch_batches': None, 'even_batches': True, 'use_seedable_sampler': True, 'non_blocking': False, 'gradient_accumulation_kwargs': None}parallelism_config: Nonedataloader_drop_last: Falsedataloader_num_workers: 0dataloader_pin_memory: Truedataloader_persistent_workers: Falsedataloader_prefetch_factor: Noneremove_unused_columns: Truelabel_names: Nonetrain_sampling_strategy: randomlength_column_name: lengthddp_find_unused_parameters: Noneddp_bucket_cap_mb: Noneddp_broadcast_buffers: Falseddp_static_graph: Noneddp_backend: Noneddp_timeout: 1800fsdp: Nonefsdp_config: Nonedeepspeed: Nonedebug: []skip_memory_metrics: Truedo_predict: Falseresume_from_checkpoint: Nonewarmup_ratio: Nonelocal_rank: -1prompts: Nonebatch_sampler: batch_samplermulti_dataset_batch_sampler: round_robinrouter_mapping: {}learning_rate_mapping: {}| Epoch | Step | Training Loss |
|---|---|---|
| 0.1895 | 500 | 0.9598 |
| 0.3791 | 1000 | 0.3102 |
| 0.5686 | 1500 | 0.2380 |
| 0.7582 | 2000 | 0.2082 |
| 0.9477 | 2500 | 0.1903 |
@inproceedings{reimers-2019-sentence-bert,
title = "Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks",
author = "Reimers, Nils and Gurevych, Iryna",
booktitle = "Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing",
month = "11",
year = "2019",
publisher = "Association for Computational Linguistics",
url = "https://arxiv.org/abs/1908.10084",
}
@misc{oord2019representationlearningcontrastivepredictive,
title={Representation Learning with Contrastive Predictive Coding},
author={Aaron van den Oord and Yazhe Li and Oriol Vinyals},
year={2019},
eprint={1807.03748},
archivePrefix={arXiv},
primaryClass={cs.LG},
url={https://arxiv.org/abs/1807.03748},
}
Base model
intfloat/multilingual-e5-small