Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks
Paper • 1908.10084 • Published • 18
How to use nikatonika/chatbot_biencoder_v2_cos_sim with sentence-transformers:
from sentence_transformers import SentenceTransformer
model = SentenceTransformer("nikatonika/chatbot_biencoder_v2_cos_sim")
sentences = [
"You haven't changed. Heard about your leg.",
"How come Every time you compliment me, it sounds like an accusation?",
"I will grant amnesty to all brothers who thrown down their arms before nightfall. And you, Ser Davos, I will allow you to travel south, a free man with a fresh horse.",
"Yeah, pulled a hamstring playing Twister. Just gonna walk it off. So, who's the girl?"
]
embeddings = model.encode(sentences)
similarities = model.similarity(embeddings, embeddings)
print(similarities.shape)
# [4, 4]This is a sentence-transformers model trained. It maps sentences & paragraphs to a 768-dimensional dense vector space and can be used for semantic textual similarity, semantic search, paraphrase mining, text classification, clustering, and more.
SentenceTransformer(
(0): Transformer({'max_seq_length': 128, 'do_lower_case': False}) with Transformer model: RobertaModel
(1): Pooling({'word_embedding_dimension': 768, 'pooling_mode_cls_token': False, 'pooling_mode_mean_tokens': True, 'pooling_mode_max_tokens': False, 'pooling_mode_mean_sqrt_len_tokens': False, 'pooling_mode_weightedmean_tokens': False, 'pooling_mode_lasttoken': False, 'include_prompt': True})
)
First install the Sentence Transformers library:
pip install -U sentence-transformers
Then you can load this model and run inference.
from sentence_transformers import SentenceTransformer
# Download from the 🤗 Hub
model = SentenceTransformer("nikatonika/chatbot_biencoder_v2_cos_sim")
# Run inference
sentences = [
"No, not yet! Takes his hands off his face. I don't know what to say.",
'Just... be straight with her.',
'Youre crazy, go crazy.',
]
embeddings = model.encode(sentences)
print(embeddings.shape)
# [3, 768]
# Get the similarity scores for the embeddings
similarities = model.similarity(embeddings, embeddings)
print(similarities.shape)
# [3, 3]
dev_evaluator and final_evaluatorTripletEvaluator| Metric | dev_evaluator | final_evaluator |
|---|---|---|
| cosine_accuracy | 0.5404 | 0.9775 |
sentence_0, sentence_1, and sentence_2| sentence_0 | sentence_1 | sentence_2 | |
|---|---|---|---|
| type | string | string | string |
| details |
|
|
|
| sentence_0 | sentence_1 | sentence_2 |
|---|---|---|
House, it's your only option. |
What do I do if my only option won't work? |
Have to think it through. Because if they say no... |
Who are you calling? |
Pizza. You like anchovies? I'm calling the Lawyer. House tucks his phone under his Parkn and starts CPR Genuine emergency. He'll okay admission. |
He wants me to lead one day. |
Suspicious. WhatEver I decide? House nods. You're setting me up. |
Laughs. Why would I do that? |
Central Perk, Chandler is getting advice from Ross and Joey. |
TripletLoss with these parameters:{
"distance_metric": "TripletDistanceMetric.EUCLIDEAN",
"triplet_margin": 5
}
per_device_train_batch_size: 64per_device_eval_batch_size: 64num_train_epochs: 30multi_dataset_batch_sampler: round_robinoverwrite_output_dir: Falsedo_predict: Falseeval_strategy: noprediction_loss_only: Trueper_device_train_batch_size: 64per_device_eval_batch_size: 64per_gpu_train_batch_size: Noneper_gpu_eval_batch_size: Nonegradient_accumulation_steps: 1eval_accumulation_steps: Nonetorch_empty_cache_steps: Nonelearning_rate: 5e-05weight_decay: 0.0adam_beta1: 0.9adam_beta2: 0.999adam_epsilon: 1e-08max_grad_norm: 1num_train_epochs: 30max_steps: -1lr_scheduler_type: linearlr_scheduler_kwargs: {}warmup_ratio: 0.0warmup_steps: 0log_level: passivelog_level_replica: warninglog_on_each_node: Truelogging_nan_inf_filter: Truesave_safetensors: Truesave_on_each_node: Falsesave_only_model: Falserestore_callback_states_from_checkpoint: Falseno_cuda: Falseuse_cpu: Falseuse_mps_device: Falseseed: 42data_seed: Nonejit_mode_eval: Falseuse_ipex: Falsebf16: Falsefp16: Falsefp16_opt_level: O1half_precision_backend: autobf16_full_eval: Falsefp16_full_eval: Falsetf32: Nonelocal_rank: 0ddp_backend: Nonetpu_num_cores: Nonetpu_metrics_debug: Falsedebug: []dataloader_drop_last: Falsedataloader_num_workers: 0dataloader_prefetch_factor: Nonepast_index: -1disable_tqdm: Falseremove_unused_columns: Truelabel_names: Noneload_best_model_at_end: Falseignore_data_skip: Falsefsdp: []fsdp_min_num_params: 0fsdp_config: {'min_num_params': 0, 'xla': False, 'xla_fsdp_v2': False, 'xla_fsdp_grad_ckpt': False}fsdp_transformer_layer_cls_to_wrap: Noneaccelerator_config: {'split_batches': False, 'dispatch_batches': None, 'even_batches': True, 'use_seedable_sampler': True, 'non_blocking': False, 'gradient_accumulation_kwargs': None}deepspeed: Nonelabel_smoothing_factor: 0.0optim: adamw_torchoptim_args: Noneadafactor: Falsegroup_by_length: Falselength_column_name: lengthddp_find_unused_parameters: Noneddp_bucket_cap_mb: Noneddp_broadcast_buffers: Falsedataloader_pin_memory: Truedataloader_persistent_workers: Falseskip_memory_metrics: Trueuse_legacy_prediction_loop: Falsepush_to_hub: Falseresume_from_checkpoint: Nonehub_model_id: Nonehub_strategy: every_savehub_private_repo: Nonehub_always_push: Falsegradient_checkpointing: Falsegradient_checkpointing_kwargs: Noneinclude_inputs_for_metrics: Falseinclude_for_metrics: []eval_do_concat_batches: Truefp16_backend: autopush_to_hub_model_id: Nonepush_to_hub_organization: Nonemp_parameters: auto_find_batch_size: Falsefull_determinism: Falsetorchdynamo: Noneray_scope: lastddp_timeout: 1800torch_compile: Falsetorch_compile_backend: Nonetorch_compile_mode: Nonedispatch_batches: Nonesplit_batches: Noneinclude_tokens_per_second: Falseinclude_num_input_tokens_seen: Falseneftune_noise_alpha: Noneoptim_target_modules: Nonebatch_eval_metrics: Falseeval_on_start: Falseuse_liger_kernel: Falseeval_use_gather_object: Falseaverage_tokens_across_devices: Falseprompts: Nonebatch_sampler: batch_samplermulti_dataset_batch_sampler: round_robin| Epoch | Step | Training Loss | dev_evaluator_cosine_accuracy | final_evaluator_cosine_accuracy |
|---|---|---|---|---|
| -1 | -1 | - | 0.5404 | - |
| 1.8315 | 500 | 1.5031 | - | - |
| 3.6630 | 1000 | 0.2306 | - | - |
| 5.4945 | 1500 | 0.0923 | - | - |
| 7.3260 | 2000 | 0.0451 | - | - |
| 9.1575 | 2500 | 0.0234 | - | - |
| 10.9890 | 3000 | 0.0133 | - | - |
| 12.8205 | 3500 | 0.0076 | - | - |
| 14.6520 | 4000 | 0.0073 | - | - |
| 16.4835 | 4500 | 0.005 | - | - |
| 18.3150 | 5000 | 0.0039 | - | - |
| 20.1465 | 5500 | 0.0038 | - | - |
| 21.9780 | 6000 | 0.0018 | - | - |
| 23.8095 | 6500 | 0.0021 | - | - |
| 25.6410 | 7000 | 0.0014 | - | - |
| 27.4725 | 7500 | 0.0009 | - | - |
| 29.3040 | 8000 | 0.0008 | - | - |
| -1 | -1 | - | - | 0.9775 |
@inproceedings{reimers-2019-sentence-bert,
title = "Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks",
author = "Reimers, Nils and Gurevych, Iryna",
booktitle = "Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing",
month = "11",
year = "2019",
publisher = "Association for Computational Linguistics",
url = "https://arxiv.org/abs/1908.10084",
}
@misc{hermans2017defense,
title={In Defense of the Triplet Loss for Person Re-Identification},
author={Alexander Hermans and Lucas Beyer and Bastian Leibe},
year={2017},
eprint={1703.07737},
archivePrefix={arXiv},
primaryClass={cs.CV}
}