SentenceTransformer based on distilbert/distilroberta-base

This is a sentence-transformers model finetuned from distilbert/distilroberta-base. It maps sentences & paragraphs to a 768-dimensional dense vector space and can be used for semantic textual similarity, semantic search, paraphrase mining, text classification, clustering, and more.

Model Details

Model Description

  • Model Type: Sentence Transformer
  • Base model: distilbert/distilroberta-base
  • Maximum Sequence Length: 128 tokens
  • Output Dimensionality: 768 dimensions
  • Similarity Function: Cosine Similarity

Model Sources

Full Model Architecture

SentenceTransformer(
  (0): Transformer({'max_seq_length': 128, 'do_lower_case': False}) with Transformer model: RobertaModel 
  (1): Pooling({'word_embedding_dimension': 768, 'pooling_mode_cls_token': False, 'pooling_mode_mean_tokens': True, 'pooling_mode_max_tokens': False, 'pooling_mode_mean_sqrt_len_tokens': False, 'pooling_mode_weightedmean_tokens': False, 'pooling_mode_lasttoken': False, 'include_prompt': True})
)

Usage

Direct Usage (Sentence Transformers)

First install the Sentence Transformers library:

pip install -U sentence-transformers

Then you can load this model and run inference.

from sentence_transformers import SentenceTransformer

# Download from the 🤗 Hub
model = SentenceTransformer("nikatonika/chatbot_sentence-transformer")
# Run inference
sentences = [
    'Like you are now? [SEP] That liver is going to somebody right now. Were doing that surgery. If you do the surgery, youll be killing a mother of four. Father of three. I was guessing. Naphthalene poisoning is the best explanation we have for whats wrong with your son. It explains the internal bleeding, the hemolytic anemia, the liver failure! it also predicts whatll happen next. If you do the surgery hes gonna lay on that table for fourteen hours while his body continues to burn fat and release poison into his system. Either way, I did you a favor. Hes awake now, youve got a chance to say goodbye.',
    'If you do the surgery hes gonna lay on that table for fourteen hours while his body continues to burn fat and release poison into',
    'I know none of that. If I did, youd be the last to know.',
]
embeddings = model.encode(sentences)
print(embeddings.shape)
# [3, 768]

# Get the similarity scores for the embeddings
similarities = model.similarity(embeddings, embeddings)
print(similarities.shape)
# [3, 3]

Evaluation

Metrics

Triplet

Metric Value
cosine_accuracy 0.9809

Training Details

Training Dataset

Unnamed Dataset

  • Size: 6,284 training samples
  • Columns: sentence_0, sentence_1, and sentence_2
  • Approximate statistics based on the first 1000 samples:
    sentence_0 sentence_1 sentence_2
    type string string string
    details
    • min: 32 tokens
    • mean: 106.25 tokens
    • max: 128 tokens
    • min: 4 tokens
    • mean: 20.84 tokens
    • max: 128 tokens
    • min: 5 tokens
    • mean: 18.96 tokens
    • max: 69 tokens
  • Samples:
    sentence_0 sentence_1 sentence_2
    I thought Everybody lied? [SEP] Told you, cant trust people. She pRobably knew she was allergic to gadolinium, figured it was an easy way to get someone to cut a hole in her throat. Cant get a picture, gonna have to get a thousand words. You actually want me to talk to the Patient? Get a history? We need to know if theres some genetic or environmental causes triggering an inflammatory response. Truth begins in lies. Think about it. Truth begins in lies. Think about it. the Krusshy and the... Krab... pizza...
    Whats that? [SEP] Her blood pressures rising. Mines rising too, course I am doing battle with a deity. In the heart, injecting the dye. Right coronary flow isnt obstructed, left coronary flow looks normal. Looks like youre wrong. Either Im right, or this test is about to go very bad. She has one... two... third ostium. How Many is she supposed to have? Dos. All the third ones doing is causing inflammation, throwing off clots, giving away the angiogram. No huMan would screw up that big! Dont worry, just one more surgery and youll be fine. She has one... two... third ostium. How Many is she supposed to have? Dos. All the third ones doing is causing inflammation, throwing off clots, giving away the angiogram. Of course I’m jokin’! I don’t take checks.
    Do me a favor!? [SEP] Mmhhmmm, I need to go peepee. Dial it up a notch and repeat. Ill be back. Ooh, girl in the boys bathroom. Very dramatic. Must be very important what you have to say to me. Yesterday your Patients tumor was 5.8 centimeters. Today its 4.6. How did that happen? At a guess, Id say Dr. House must be really really good ì why am I wasting him on hiccups?ù I wash before and after. You also requisitioned 20cc of ethanol what Patient was that for? Or are you planning a party? I was gonna say leave,ù but that works. I was gonna say leave,ù but that works. I seem to recall them giving you a bit of trouble as well.
  • Loss: TripletLoss with these parameters:
    {
        "distance_metric": "TripletDistanceMetric.EUCLIDEAN",
        "triplet_margin": 5
    }
    

Training Hyperparameters

Non-Default Hyperparameters

  • eval_strategy: steps
  • num_train_epochs: 1
  • multi_dataset_batch_sampler: round_robin

All Hyperparameters

Click to expand
  • overwrite_output_dir: False
  • do_predict: False
  • eval_strategy: steps
  • prediction_loss_only: True
  • per_device_train_batch_size: 8
  • per_device_eval_batch_size: 8
  • per_gpu_train_batch_size: None
  • per_gpu_eval_batch_size: None
  • gradient_accumulation_steps: 1
  • eval_accumulation_steps: None
  • torch_empty_cache_steps: None
  • learning_rate: 5e-05
  • weight_decay: 0.0
  • adam_beta1: 0.9
  • adam_beta2: 0.999
  • adam_epsilon: 1e-08
  • max_grad_norm: 1
  • num_train_epochs: 1
  • max_steps: -1
  • lr_scheduler_type: linear
  • lr_scheduler_kwargs: {}
  • warmup_ratio: 0.0
  • warmup_steps: 0
  • log_level: passive
  • log_level_replica: warning
  • log_on_each_node: True
  • logging_nan_inf_filter: True
  • save_safetensors: True
  • save_on_each_node: False
  • save_only_model: False
  • restore_callback_states_from_checkpoint: False
  • no_cuda: False
  • use_cpu: False
  • use_mps_device: False
  • seed: 42
  • data_seed: None
  • jit_mode_eval: False
  • use_ipex: False
  • bf16: False
  • fp16: False
  • fp16_opt_level: O1
  • half_precision_backend: auto
  • bf16_full_eval: False
  • fp16_full_eval: False
  • tf32: None
  • local_rank: 0
  • ddp_backend: None
  • tpu_num_cores: None
  • tpu_metrics_debug: False
  • debug: []
  • dataloader_drop_last: False
  • dataloader_num_workers: 0
  • dataloader_prefetch_factor: None
  • past_index: -1
  • disable_tqdm: False
  • remove_unused_columns: True
  • label_names: None
  • load_best_model_at_end: False
  • ignore_data_skip: False
  • fsdp: []
  • fsdp_min_num_params: 0
  • fsdp_config: {'min_num_params': 0, 'xla': False, 'xla_fsdp_v2': False, 'xla_fsdp_grad_ckpt': False}
  • fsdp_transformer_layer_cls_to_wrap: None
  • accelerator_config: {'split_batches': False, 'dispatch_batches': None, 'even_batches': True, 'use_seedable_sampler': True, 'non_blocking': False, 'gradient_accumulation_kwargs': None}
  • deepspeed: None
  • label_smoothing_factor: 0.0
  • optim: adamw_torch
  • optim_args: None
  • adafactor: False
  • group_by_length: False
  • length_column_name: length
  • ddp_find_unused_parameters: None
  • ddp_bucket_cap_mb: None
  • ddp_broadcast_buffers: False
  • dataloader_pin_memory: True
  • dataloader_persistent_workers: False
  • skip_memory_metrics: True
  • use_legacy_prediction_loop: False
  • push_to_hub: False
  • resume_from_checkpoint: None
  • hub_model_id: None
  • hub_strategy: every_save
  • hub_private_repo: None
  • hub_always_push: False
  • gradient_checkpointing: False
  • gradient_checkpointing_kwargs: None
  • include_inputs_for_metrics: False
  • include_for_metrics: []
  • eval_do_concat_batches: True
  • fp16_backend: auto
  • push_to_hub_model_id: None
  • push_to_hub_organization: None
  • mp_parameters:
  • auto_find_batch_size: False
  • full_determinism: False
  • torchdynamo: None
  • ray_scope: last
  • ddp_timeout: 1800
  • torch_compile: False
  • torch_compile_backend: None
  • torch_compile_mode: None
  • dispatch_batches: None
  • split_batches: None
  • include_tokens_per_second: False
  • include_num_input_tokens_seen: False
  • neftune_noise_alpha: None
  • optim_target_modules: None
  • batch_eval_metrics: False
  • eval_on_start: False
  • use_liger_kernel: False
  • eval_use_gather_object: False
  • average_tokens_across_devices: False
  • prompts: None
  • batch_sampler: batch_sampler
  • multi_dataset_batch_sampler: round_robin

Training Logs

Epoch Step Training Loss dev_evaluator_cosine_accuracy
-1 -1 - 0.7078
0.2545 200 - 0.9255
0.5089 400 - 0.9701
0.6361 500 1.6621 -
0.7634 600 - 0.9752
1.0 786 - 0.9790
-1 -1 - 0.9790
0.2545 200 - 0.9752
0.5089 400 - 0.9790
0.6361 500 0.298 -
0.7634 600 - 0.9790
1.0 786 - 0.9803
-1 -1 - 0.9803
0.2545 200 - 0.9777
0.5089 400 - 0.9796
0.6361 500 0.0783 -
0.7634 600 - 0.9809

Framework Versions

  • Python: 3.11.11
  • Sentence Transformers: 3.4.1
  • Transformers: 4.49.0
  • PyTorch: 2.6.0+cu124
  • Accelerate: 1.3.0
  • Datasets: 3.3.2
  • Tokenizers: 0.21.0

Citation

BibTeX

Sentence Transformers

@inproceedings{reimers-2019-sentence-bert,
    title = "Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks",
    author = "Reimers, Nils and Gurevych, Iryna",
    booktitle = "Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing",
    month = "11",
    year = "2019",
    publisher = "Association for Computational Linguistics",
    url = "https://arxiv.org/abs/1908.10084",
}

TripletLoss

@misc{hermans2017defense,
    title={In Defense of the Triplet Loss for Person Re-Identification},
    author={Alexander Hermans and Lucas Beyer and Bastian Leibe},
    year={2017},
    eprint={1703.07737},
    archivePrefix={arXiv},
    primaryClass={cs.CV}
}
Downloads last month
48
Safetensors
Model size
82.1M params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for nikatonika/chatbot_sentence-transformer

Finetuned
(793)
this model
Quantizations
1 model

Papers for nikatonika/chatbot_sentence-transformer

Evaluation results