zakerzel/upy-u2t02-simcse-supervised

Sentence embedding model trained for UPY Trends in Data Science U2T02 using SimCSE and bert-base-uncased.

Training recipe

  • Mode: supervised
  • Base encoder: bert-base-uncased
  • Training data: supplied snli_train_100k.jsonl
  • Training examples: 33,351
  • Batch size: 64
  • Learning rate: 5e-05
  • Epochs: 3
  • Temperature: 0.05
  • Dropout: 0.1
  • Pooling: cls
  • Max sequence length: 64
  • Seed: 42
  • MLP policy: train-only projection; discarded for evaluation/export
  • Hard negatives: 9,488 of 33,351 supervised pairs contained a real contradiction hard negative.

Evaluation

The checkpoint was selected using STS-B dev Spearman only.

  • STS-B dev Spearman x100: 82.29
  • STS-B test Spearman x100: 78.32
  • Reloaded/exported test Spearman x100: 78.32
  • Test alignment: 0.1377
  • Test uniformity: -2.2771

Evaluation uses L2-normalized sentence embeddings, cosine similarity and Spearman correlation with STS-B human scores. No regressor is used.

Intended use

Educational sentence-similarity and retrieval experiments in English.

Limitations

  • The model was trained on a 100k-record SNLI subset rather than the full data used in the original SimCSE paper.
  • SNLI is dominated by image-caption style sentences, so domain coverage is limited.
  • This model is English-only and was not evaluated for multilingual or specialized-domain use.

Reference

Gao, T., Yao, X., & Chen, D. (2021). SimCSE: Simple Contrastive Learning of Sentence Embeddings. EMNLP 2021.

Downloads last month
15
Safetensors
Model size
0.1B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support