SimCSE (supervised) trained from bert-base-uncased on SNLI

A sentence embedding model that replicates the SimCSE training method (Gao, Yao and Chen, 2021) for Universidad Politécnica de Yucatán project. It maps sentences to 768-dimensional vectors that you compare with cosine similarity.

Usage

from sentence_transformers import SentenceTransformer
model = SentenceTransformer("ArturoBE21/simcse-bert-base-snli-sup")
emb = model.encode(["A man is playing a guitar.", "Someone plays an instrument."])

Training data

33,351 (premise, entailment hypothesis) pairs from a 100k SNLI subset; contradiction hypotheses as hard negatives for about 28% of the pairs. A specific SNLI subset was used, which is much smaller than the data in the paper.

Recipe

  • Start: google-bert/bert-base-uncased, all weights trained
  • Objective: contrastive loss with entailment pairs as positives and contradictions as hard negatives
  • Pooling: [CLS] with an MLP head during training; the MLP is kept at inference
  • Temperature 0.05, dropout 0.1, learning rate 3e-05, batch size 128, epochs 3, max length 32
  • Seed 44, precision float16, hardware Tesla T4
  • Checkpoint chosen by best Spearman on the STS-B dev set, evaluated every 50 steps (best at step 100)

Results (STS-B, Spearman x 100, cosine similarity, no regressor)

Split Spearman
dev 82.29
test 78.26

Alignment 0.169 and uniformity -3.091 on STS-B dev (lower is better).

For reference, the paper reports 86.2 dev and 84.25 test for its supervised BERT-base model, trained on different data. The reasons for the difference are analyzed in the project report.

Limitations

  • Trained on SNLI, which is made of short English image captions, so it works best on simple everyday sentences.
  • Only evaluated on STS-B. Do not assume it is good for other domains, long documents or other languages.
  • Sentences are truncated at 128 tokens.
  • It inherits the biases of BERT and of the crowd-written SNLI sentences.
  • Results are from the seed with the best STS-B dev score among 3 seeds (42, 43, 44). Scores vary across seeds; all runs are in training_logs/.
Downloads last month
50
Safetensors
Model size
0.1B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for ArturoBE21/simcse-bert-base-snli-sup

Finetuned
(7010)
this model

Dataset used to train ArturoBE21/simcse-bert-base-snli-sup