StanfordSCALE/assertions_llm_annotated_talkmoves
Viewer • Updated • 6.43k • 16
How to use StanfordSCALE/assertion_sentence_has_number with setfit:
from setfit import SetFitModel
model = SetFitModel.from_pretrained("StanfordSCALE/assertion_sentence_has_number")This classifier was trained for EduBehaviors: Assertion-based schemas for auditable dialogue coding and is usable through the Python package EduBehaviors-kit. This classifier was trained on an LLM-annotated subset of teacher utterances from the TalkMoves Dataset. See the Datasets section below for more information.
| Dataset | Split | Size |
|---|---|---|
| StanfordSCALE/assertions_llm_annotated_talkmoves | train | 3,430 (53.3%) |
| StanfordSCALE/assertions_llm_annotated_talkmoves | dev | 858 (13.3%) |
| StanfordSCALE/assertions_llm_annotated_talkmoves | test | 2,146 (33.4%) |
This model's columns are assertion_sentence_has_number and split_sentence_has_number.
Base rate (share of rows labeled as True): 25.5% overall — 24.4% train, 28.1% dev, 26.1% test.
Labels were generated with LLM annotators. Krippendorff's alpha for this assertion is 0.755.
| Parameter | Value |
|---|---|
| Base model (body) | sentence-transformers/paraphrase-mpnet-base-v2 |
| Head | LogisticRegression |
| Body learning rate | 2e-05 |
| Head learning rate | 0.01 |
| Batch size | 16 (contrastive phase) / 32 (head) |
| Epochs | 10 |
| Max steps | 5000 (contrastive phase) |
| Eval max steps | 100 |
| Seed | 20260904 |
| Mixed precision | enabled on GPU |
| Split | n | Base rate | Precision | Recall | F1 (positive class) | ROC-AUC | Average precision |
|---|---|---|---|---|---|---|---|
| dev | 858 | 28.1% | 0.881 | 0.892 | 0.887 | 0.982 | 0.942 |
| test | 2,146 | 26.1% | 0.868 | 0.880 | 0.874 | 0.983 | 0.949 |
The model was trained on text built as:
{utterance}
The utterance is passed through as-is.
pip install setfit
from setfit import SetFitModel
model = SetFitModel.from_pretrained("StanfordSCALE/assertion_sentence_has_number")
text = 'Take 30 seconds talk to your group and then were going to come back and kind of put up all the words we think of when we think of modeling'
model.predict([text]) # -> array([1]) when the assertion holds
model.predict_proba([text]) # -> [[P(no), P(yes)]]
@misc{assertion_sentence_has_number,
author = {Stanford SCALE Initiative},
title = {Assertion classifier: sentence has number},
year = {2026},
url = {https://huggingface.co/StanfordSCALE/assertion_sentence_has_number}
}