Instructions to use Perry-DLC/upy-u2t02-simcse with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- sentence-transformers
How to use Perry-DLC/upy-u2t02-simcse with sentence-transformers:
from sentence_transformers import SentenceTransformer model = SentenceTransformer("Perry-DLC/upy-u2t02-simcse") sentences = [ "That is a happy person", "That is a happy dog", "That is a very happy person", "Today is a sunny day" ] embeddings = model.encode(sentences) similarities = model.similarity(embeddings, embeddings) print(similarities.shape) # [4, 4] - Notebooks
- Google Colab
- Kaggle
U2T02 SimCSE sentence embeddings
Team
Karen Cardiel Olea (2209039), Jorge Ramiro Chay Koyoc (2309052), Angeles Alejandra Cruz Legorreta (2309064), Diego Jesús Loría Campos (2309140), Brad Robles García (2309198).
Universidad Politécnica de Yucatán. Trends in Data Science. Professor: Dexter Enrique Gómez Ek.
Model and selection
Unsupervised SimCSE, trained from bert-base-uncased. This repository contains unsup_s42, selected by STS-B dev Spearman from the primary seed-42 unsupervised/supervised pair. It is not the maximum-dev model across all repeated seeds: unsup_s43 obtained 77.7050 dev, compared with 77.1732 for the published seed-42 model. Test was not used to change the published choice.
Output: 768-dimensional normalized vectors. Inference uses CLS before BERT's dense/tanh pooler. Maximum sequence length: 64 tokens.
Training data
Provided SNLI train-only 100,000-record subset, with 165,529 distinct premise/hypothesis sentences. This unsupervised model uses those sentences without labels. No STS-B training examples, Wikipedia, MNLI or extra corpus were used.
Input file SHA-256: 0927a0b1d2a447b9236fd9699d6a37a4a3542c5def178d701de1c38827997f88.
Recipe
Seed 42; batch 64; one epoch; learning rate 3e-5; AdamW with weight decay 0; linear schedule with 10% warmup; gradient clipping norm 1; temperature 0.05; hidden/attention dropout 0.1; independent masks for the two views. Training uses CLS with BERT dense/tanh pooler. Dev is checked every 250 steps and at the final step; the highest-dev checkpoint is saved (step 1500 for this run). NVIDIA L4 in Google Colab, PyTorch 2.11.0+cu130, Transformers 5.17.0, Sentence Transformers 6.1.0, Datasets 5.0.1.
Evaluation
Dataset: sentence-transformers/stsb, dev 1,500 pairs and test 1,379 pairs. Independently embed both sentences, normalize, use cosine similarity, and compute Spearman without a regressor. Spearman is multiplied by 100.
| Metric | Original checkpoint | Hub reload |
|---|---|---|
| Dev Spearman x100 | 77.173179 | not retuned |
| Test Spearman x100 | 69.841386 | 69.841380 |
| Test alignment | 0.227365896 | 0.227365896 |
| Test uniformity | -2.376543999 | -2.376543999 |
Alignment is mean squared distance for normalized STS-B test pairs with human score >=4/5. Uniformity uses t=2, 20,000 sampled distinct-index vector pairs, RNG seed 123. Duplicate text can exist at different indices. The test pass after Hub reload is the required export verification, not model selection.
Usage
from sentence_transformers import SentenceTransformer
model = SentenceTransformer('Perry-DLC/upy-u2t02-simcse')
vectors = model.encode(
['A man with a hard hat is dancing.',
'A man wearing a hard hat is dancing.'],
normalize_embeddings=True
)
print(float(vectors[0] @ vectors[1]))
Limitations
English NLI domain bias; inputs truncated at 64 tokens; no validation for long documents, other languages or specialized domains. Cosine is not a calibrated human rating or proof of factual identity. The model assigns high similarity to headlines with different locations and magnitudes but the same earthquake template. Only three training seeds were investigated in the accompanying report. This model does not reproduce the paper's full-data results.
References
- Gao, Yao and Chen (2021): https://aclanthology.org/2021.emnlp-main.552/
- Wang and Isola (2020): https://proceedings.mlr.press/v119/wang20k.html
- SNLI: https://huggingface.co/datasets/stanfordnlp/snli
- STS-B: https://huggingface.co/datasets/sentence-transformers/stsb
- Downloads last month
- 13
Model tree for Perry-DLC/upy-u2t02-simcse
Base model
google-bert/bert-base-uncased