U2T02 SimCSE sentence embeddings

Team

Karen Cardiel Olea (2209039), Jorge Ramiro Chay Koyoc (2309052), Angeles Alejandra Cruz Legorreta (2309064), Diego Jesús Loría Campos (2309140), Brad Robles García (2309198).

Universidad Politécnica de Yucatán. Trends in Data Science. Professor: Dexter Enrique Gómez Ek.

Model and selection

Unsupervised SimCSE, trained from bert-base-uncased. This repository contains unsup_s42, selected by STS-B dev Spearman from the primary seed-42 unsupervised/supervised pair. It is not the maximum-dev model across all repeated seeds: unsup_s43 obtained 77.7050 dev, compared with 77.1732 for the published seed-42 model. Test was not used to change the published choice.

Output: 768-dimensional normalized vectors. Inference uses CLS before BERT's dense/tanh pooler. Maximum sequence length: 64 tokens.

Training data

Provided SNLI train-only 100,000-record subset, with 165,529 distinct premise/hypothesis sentences. This unsupervised model uses those sentences without labels. No STS-B training examples, Wikipedia, MNLI or extra corpus were used.

Input file SHA-256: 0927a0b1d2a447b9236fd9699d6a37a4a3542c5def178d701de1c38827997f88.

Recipe

Seed 42; batch 64; one epoch; learning rate 3e-5; AdamW with weight decay 0; linear schedule with 10% warmup; gradient clipping norm 1; temperature 0.05; hidden/attention dropout 0.1; independent masks for the two views. Training uses CLS with BERT dense/tanh pooler. Dev is checked every 250 steps and at the final step; the highest-dev checkpoint is saved (step 1500 for this run). NVIDIA L4 in Google Colab, PyTorch 2.11.0+cu130, Transformers 5.17.0, Sentence Transformers 6.1.0, Datasets 5.0.1.

Evaluation

Dataset: sentence-transformers/stsb, dev 1,500 pairs and test 1,379 pairs. Independently embed both sentences, normalize, use cosine similarity, and compute Spearman without a regressor. Spearman is multiplied by 100.

Metric Original checkpoint Hub reload
Dev Spearman x100 77.173179 not retuned
Test Spearman x100 69.841386 69.841380
Test alignment 0.227365896 0.227365896
Test uniformity -2.376543999 -2.376543999

Alignment is mean squared distance for normalized STS-B test pairs with human score >=4/5. Uniformity uses t=2, 20,000 sampled distinct-index vector pairs, RNG seed 123. Duplicate text can exist at different indices. The test pass after Hub reload is the required export verification, not model selection.

Usage

from sentence_transformers import SentenceTransformer
model = SentenceTransformer('Perry-DLC/upy-u2t02-simcse')
vectors = model.encode(
    ['A man with a hard hat is dancing.',
     'A man wearing a hard hat is dancing.'],
    normalize_embeddings=True
)
print(float(vectors[0] @ vectors[1]))

Limitations

English NLI domain bias; inputs truncated at 64 tokens; no validation for long documents, other languages or specialized domains. Cosine is not a calibrated human rating or proof of factual identity. The model assigns high similarity to headlines with different locations and magnitudes but the same earthquake template. Only three training seeds were investigated in the accompanying report. This model does not reproduce the paper's full-data results.

References

Downloads last month
13
Safetensors
Model size
0.1B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Perry-DLC/upy-u2t02-simcse

Finetuned
(7014)
this model