SentenceTransformer based on google-bert/bert-base-uncased

This is a sentence-transformers model finetuned from google-bert/bert-base-uncased on the wiki1m-for-simcse dataset. It maps sentences & paragraphs to a 768-dimensional dense vector space and can be used for retrieval.

Model Details

Model Description

  • Model Type: Sentence Transformer
  • Base model: google-bert/bert-base-uncased
  • Maximum Sequence Length: 75 tokens
  • Output Dimensionality: 768 dimensions
  • Similarity Function: Cosine Similarity
  • Supported Modality: Text
  • Training Dataset:
    • wiki1m-for-simcse

Model Sources

Full Model Architecture

SentenceTransformer(
  (0): Transformer({'transformer_task': 'feature-extraction', 'modality_config': {'text': {'method': 'forward', 'method_output_name': 'last_hidden_state'}}, 'module_output_name': 'token_embeddings', 'architecture': 'BertModel'})
  (1): Pooling({'embedding_dimension': 768, 'pooling_mode': 'cls', 'include_prompt': True})
)

Usage

Direct Usage (Sentence Transformers)

First install the Sentence Transformers library:

pip install -U sentence-transformers

Then you can load this model and run inference.

from sentence_transformers import SentenceTransformer

# Download from the 🤗 Hub
model = SentenceTransformer("kwondw/bert-base-uncased-tsdae")
# Run inference
sentences = [
    'uses a high-frequency, encodes the audio and can be distributed over the waves generating a bridge between the analog and digital.',
    'This app uses a high-frequency algorithm, which encodes the audio information and can be distributed over the radio waves, generating a bridge between the analog and digital world.',
    'The book has been published by many organizations around the world:',
]
embeddings = model.encode(sentences)
print(embeddings.shape)
# [3, 768]

# Get the similarity scores for the embeddings
similarities = model.similarity(embeddings, embeddings)
print(similarities)
# tensor([[1.0000, 0.9122, 0.4017],
#         [0.9122, 1.0000, 0.4779],
#         [0.4017, 0.4779, 1.0000]])

Evaluation

Metrics

Semantic Similarity

Metric sts-dev sts-test
pearson_cosine 0.6503 0.6159
spearman_cosine 0.6557 0.6199

Training Details

Training Dataset

wiki1m-for-simcse

  • Dataset: wiki1m-for-simcse
  • Size: 990,000 training samples
  • Columns: noisy and text
  • Approximate statistics based on the first 1000 samples:
    noisy text
    type string string
    details
    • min: 3 tokens
    • mean: 17.73 tokens
    • max: 61 tokens
    • min: 3 tokens
    • mean: 28.09 tokens
    • max: 75 tokens
  • Samples:
    noisy text
    Zagreb together the Army It liberated Zagreb on May 8, together with parts of the 2nd Army.
    was curious to learn about fascism from the source however in 1933 an trip met wife at London home to Campbell was curious to learn about fascism from the source however, so in 1933 during an overseas business trip, he met with Sir Oswald Mosley and wife Lady Cynthia at their London home to discuss the matter.
    Republican Henry H. Crapo defeated Democratic nominee William H. Fenton with 55.15 of the Republican nominee Henry H. Crapo defeated Democratic nominee William H. Fenton with 55.15% of the vote.
  • Loss: DenoisingAutoEncoderLoss with these parameters:
    {
        "decoder_name_or_path": "google-bert/bert-base-uncased",
        "need_retokenization": false
    }
    

Evaluation Dataset

wiki1m-for-simcse

  • Dataset: wiki1m-for-simcse
  • Size: 10,000 evaluation samples
  • Columns: noisy and text
  • Approximate statistics based on the first 1000 samples:
    noisy text
    type string string
    details
    • min: 3 tokens
    • mean: 17.48 tokens
    • max: 64 tokens
    • min: 3 tokens
    • mean: 27.57 tokens
    • max: 75 tokens
  • Samples:
    noisy text
    Burnham are. Burnham said, "They are hypocrites.
    Jocano further emphasizes advancements in the ceramic industry, which led in trade and the of jar in Philippines Jocano further emphasizes the advancements made in the ceramic industry, which led to progress in trade and the eventual use of jar burials in the Philippines.
    regarded with severity especially harmful in current generation the generation of freedom and") since strict might lead individuals not to comply with the. Yosef regarded ruling with severity as especially harmful in the current generation ("the generation of freedom and liberty"), since strict ruling might lead individuals not to comply with the Halakha.
  • Loss: DenoisingAutoEncoderLoss with these parameters:
    {
        "decoder_name_or_path": "google-bert/bert-base-uncased",
        "need_retokenization": false
    }
    

Training Hyperparameters

Non-Default Hyperparameters

  • per_device_train_batch_size: 16
  • per_device_eval_batch_size: 16
  • learning_rate: 3e-05
  • num_train_epochs: 1
  • warmup_ratio: 0.1
  • fp16: True

All Hyperparameters

Click to expand
  • overwrite_output_dir: False
  • do_predict: False
  • prediction_loss_only: True
  • per_device_train_batch_size: 16
  • per_device_eval_batch_size: 16
  • per_gpu_train_batch_size: None
  • per_gpu_eval_batch_size: None
  • gradient_accumulation_steps: 1
  • eval_accumulation_steps: None
  • torch_empty_cache_steps: None
  • learning_rate: 3e-05
  • weight_decay: 0.0
  • adam_beta1: 0.9
  • adam_beta2: 0.999
  • adam_epsilon: 1e-08
  • max_grad_norm: 1.0
  • num_train_epochs: 1
  • max_steps: -1
  • lr_scheduler_type: linear
  • lr_scheduler_kwargs: None
  • warmup_ratio: 0.1
  • warmup_steps: 0
  • log_level: passive
  • log_level_replica: warning
  • log_on_each_node: True
  • logging_nan_inf_filter: True
  • save_safetensors: True
  • save_on_each_node: False
  • save_only_model: False
  • restore_callback_states_from_checkpoint: False
  • no_cuda: False
  • use_cpu: False
  • use_mps_device: False
  • seed: 42
  • data_seed: None
  • jit_mode_eval: False
  • bf16: False
  • fp16: True
  • fp16_opt_level: O1
  • half_precision_backend: auto
  • bf16_full_eval: False
  • fp16_full_eval: False
  • tf32: None
  • local_rank: 0
  • ddp_backend: None
  • tpu_num_cores: None
  • tpu_metrics_debug: False
  • debug: []
  • dataloader_drop_last: True
  • dataloader_num_workers: 0
  • dataloader_prefetch_factor: None
  • past_index: -1
  • disable_tqdm: False
  • remove_unused_columns: True
  • label_names: None
  • load_best_model_at_end: False
  • ignore_data_skip: False
  • fsdp: []
  • fsdp_min_num_params: 0
  • fsdp_config: {'min_num_params': 0, 'xla': False, 'xla_fsdp_v2': False, 'xla_fsdp_grad_ckpt': False}
  • fsdp_transformer_layer_cls_to_wrap: None
  • accelerator_config: {'split_batches': False, 'dispatch_batches': None, 'even_batches': True, 'use_seedable_sampler': True, 'non_blocking': False, 'gradient_accumulation_kwargs': None}
  • parallelism_config: None
  • deepspeed: None
  • label_smoothing_factor: 0.0
  • optim: adamw_torch_fused
  • optim_args: None
  • adafactor: False
  • group_by_length: False
  • length_column_name: length
  • project: huggingface
  • trackio_space_id: trackio
  • ddp_find_unused_parameters: None
  • ddp_bucket_cap_mb: None
  • ddp_broadcast_buffers: False
  • dataloader_pin_memory: True
  • dataloader_persistent_workers: False
  • skip_memory_metrics: True
  • use_legacy_prediction_loop: False
  • push_to_hub: False
  • resume_from_checkpoint: None
  • hub_model_id: None
  • hub_strategy: every_save
  • hub_private_repo: None
  • hub_always_push: False
  • hub_revision: None
  • gradient_checkpointing: False
  • gradient_checkpointing_kwargs: None
  • include_inputs_for_metrics: False
  • include_for_metrics: []
  • eval_do_concat_batches: True
  • fp16_backend: auto
  • push_to_hub_model_id: None
  • push_to_hub_organization: None
  • mp_parameters:
  • auto_find_batch_size: False
  • full_determinism: False
  • torchdynamo: None
  • ray_scope: last
  • ddp_timeout: 1800
  • torch_compile: False
  • torch_compile_backend: None
  • torch_compile_mode: None
  • include_tokens_per_second: False
  • include_num_input_tokens_seen: no
  • neftune_noise_alpha: None
  • optim_target_modules: None
  • batch_eval_metrics: False
  • eval_on_start: False
  • use_liger_kernel: False
  • liger_kernel_config: None
  • eval_use_gather_object: False
  • average_tokens_across_devices: True
  • prompts: None
  • batch_sampler: batch_sampler
  • multi_dataset_batch_sampler: proportional
  • router_mapping: {}
  • learning_rate_mapping: {}

Training Logs

Epoch Step Training Loss Validation Loss sts-dev_spearman_cosine sts-test_spearman_cosine
-1 -1 - - 0.3173 -
0.0323 1000 5.932 - - -
0.0646 2000 4.3083 - - -
0.0970 3000 3.8935 - - -
0.1000 3094 - 3.6280 0.7205 -
0.1293 4000 3.5923 - - -
0.1616 5000 3.4069 - - -
0.1939 6000 3.2878 - - -
0.2000 6188 - 3.1559 0.6996 -
0.2263 7000 3.1971 - - -
0.2586 8000 3.1312 - - -
0.2909 9000 3.0869 - - -
0.3000 9282 - 2.9709 0.6858 -
0.3232 10000 3.0341 - - -
0.3556 11000 2.9983 - - -
0.3879 12000 2.9585 - - -
0.4000 12376 - 2.8571 0.6755 -
0.4202 13000 2.9275 - - -
0.4525 14000 2.9047 - - -
0.4849 15000 2.8768 - - -
0.5000 15470 - 2.7661 0.6758 -
0.5172 16000 2.853 - - -
0.5495 17000 2.8265 - - -
0.5818 18000 2.8025 - - -
0.6001 18564 - 2.7021 0.6665 -
0.6142 19000 2.8065 - - -
0.6465 20000 2.7683 - - -
0.6788 21000 2.75 - - -
0.7001 21658 - 2.6579 0.6524 -
0.7111 22000 2.7425 - - -
0.7434 23000 2.7328 - - -
0.7758 24000 2.7114 - - -
0.8001 24752 - 2.6154 0.6605 -
0.8081 25000 2.6982 - - -
0.8404 26000 2.6898 - - -
0.8727 27000 2.6775 - - -
0.9001 27846 - 2.5901 0.6557 -
0.9051 28000 2.6655 - - -
0.9374 29000 2.6682 - - -
0.9697 30000 2.6622 - - -
-1 -1 - - - 0.6199

Training Time

  • Training: 2.8 hours
  • Evaluation: 5.9 minutes
  • Total: 2.9 hours

Framework Versions

  • Python: 3.12.13
  • Sentence Transformers: 5.4.1
  • Transformers: 4.57.6
  • PyTorch: 2.10.0+cu128
  • Accelerate: 1.13.0
  • Datasets: 5.0.0
  • Tokenizers: 0.22.2

Citation

BibTeX

Sentence Transformers

@inproceedings{reimers-2019-sentence-bert,
    title = "Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks",
    author = "Reimers, Nils and Gurevych, Iryna",
    booktitle = "Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing",
    month = "11",
    year = "2019",
    publisher = "Association for Computational Linguistics",
    url = "https://arxiv.org/abs/1908.10084",
}

DenoisingAutoEncoderLoss

@inproceedings{wang-2021-TSDAE,
    title = "TSDAE: Using Transformer-based Sequential Denoising Auto-Encoderfor Unsupervised Sentence Embedding Learning",
    author = "Wang, Kexin and Reimers, Nils and Gurevych, Iryna",
    booktitle = "Findings of the Association for Computational Linguistics: EMNLP 2021",
    month = nov,
    year = "2021",
    address = "Punta Cana, Dominican Republic",
    publisher = "Association for Computational Linguistics",
    pages = "671--688",
    url = "https://arxiv.org/abs/2104.06979",
}
Downloads last month
-
Safetensors
Model size
0.1B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for kwondw/bert-base-uncased-tsdae

Finetuned
(6863)
this model

Papers for kwondw/bert-base-uncased-tsdae

Evaluation results