Assertion: sentence has number

This classifier was trained for EduBehaviors: Assertion-based schemas for auditable dialogue coding and is usable through the Python package EduBehaviors-kit. This classifier was trained on an LLM-annotated subset of teacher utterances from the TalkMoves Dataset. See the Datasets section below for more information.


Training Details

Datasets

This model's columns are assertion_sentence_has_number and split_sentence_has_number.

Base rate (share of rows labeled as True): 25.5% overall — 24.4% train, 28.1% dev, 26.1% test.

Labels and annotation

Labels were generated with LLM annotators. Krippendorff's alpha for this assertion is 0.755.

Hyperparameters

Parameter Value
Base model (body) sentence-transformers/paraphrase-mpnet-base-v2
Head LogisticRegression
Body learning rate 2e-05
Head learning rate 0.01
Batch size 16 (contrastive phase) / 32 (head)
Epochs 10
Max steps 5000 (contrastive phase)
Eval max steps 100
Seed 20260904
Mixed precision enabled on GPU

Evaluation

Results

Split n Base rate Precision Recall F1 (positive class) ROC-AUC Average precision
dev 858 28.1% 0.881 0.892 0.887 0.982 0.942
test 2,146 26.1% 0.868 0.880 0.874 0.983 0.949

Limitations

  • Labels come from LLM annotators, not human coders. Agreement between annotators with Krippendorff's Alpha is 0.755.
  • Trained on teacher utterances only. Behaviour on student speech is untested.

How to Use

Message Structure

The model was trained on text built as:

{utterance}

The utterance is passed through as-is.

Running instructions

pip install setfit
from setfit import SetFitModel

model = SetFitModel.from_pretrained("StanfordSCALE/assertion_sentence_has_number")

text = 'Take 30 seconds talk to your group and then were going to come back and kind of put up all the words we think of when we think of modeling'
model.predict([text])        # -> array([1]) when the assertion holds
model.predict_proba([text])  # -> [[P(no), P(yes)]]

Citation

@misc{assertion_sentence_has_number,
  author = {Stanford SCALE Initiative},
  title  = {Assertion classifier: sentence has number},
  year   = {2026},
  url    = {https://huggingface.co/StanfordSCALE/assertion_sentence_has_number}
}
Downloads last month
15
Safetensors
Model size
0.1B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for StanfordSCALE/assertion_sentence_has_number

Finetuned
(372)
this model

Dataset used to train StanfordSCALE/assertion_sentence_has_number

Collection including StanfordSCALE/assertion_sentence_has_number

Evaluation results