qlm-nto-classifier

A DeBERTa-v3 classifier of pedagogical talk moves in tutoring and classroom discourse, built and deployed at Quantum Learning Machines before June 2026. Where its sibling model (qlm-map-classifier) reads the student's side of an exchange, this model reads the teacher's or tutor's side: it classifies instructional turns into pedagogical move categories — Socratic questioning, revoicing, scaffolding, redirection, worked-example presentation, praise types, direct instruction, and related moves — so that tutoring quality can be evaluated as behavior, not vibes.

What it does in production

The model supports two production surfaces. In Tutor Lab, QLM's evaluation infrastructure for AI tutoring models, talk-move classification contributes to the pedagogical-quality dimension of tutor evaluation — identifying whether a tutor's turns exhibit the moves the tutoring literature associates with learning, rather than merely fluent text. In TeachProof's discourse-analysis engine, the same move taxonomy underlies per-turn analysis of practice sessions and uploaded lesson transcripts (questioning patterns, revoicing counts, praise specificity), feeding research-cited coaching aligned to established observation frameworks. In both surfaces, classifier outputs are aggregated and reviewed — no single classification drives a consequential judgment.

Evaluation

We have not yet published formal benchmark figures for this model, and we won't imply any here: per our reporting policy, no number appears in our materials without a versioned result file behind it. Quantitative evaluation — including inter-judge agreement with human raters and reliability-filtered comparison data — is being produced through the Tutor Lab pipeline (Bradley-Terry scoring, Dawid-Skene rater reliability, transitivity auditing) and will be added to this card when it exists. Until then, treat this as a production-deployed component published for transparency and research use, not a benchmarked reference model.

Known limitations

Talk-move boundaries are genuinely fuzzy: a single turn often performs multiple moves, and single-label classification flattens that. The model is trained on English mathematics-tutoring discourse; transfer to other subjects, registers (whole-class vs. one-on-one), and languages is untested. Move presence is not move quality — detecting a Socratic question says nothing about whether it was a good one; quality judgments in our systems come from separate evaluation layers with human raters.

Intended use

  • Research on pedagogical discourse, tutoring-move analysis, and AI-tutor evaluation
  • As a component in aggregate, human-reviewed evaluation pipelines
  • Extending or critiquing move-taxonomy approaches to tutoring quality

Out-of-scope use

  • Consequential evaluation of individual human teachers — hiring, retention, compensation, or formal performance ratings. This model is not validated for stakes, and talk-move counts are not teaching quality.
  • Standalone automated scoring of tutors or teachers without human raters in the loop
  • Non-English or non-instructional discourse

Training data

Instructional-discourse data from QLM's production tutoring platforms and expert-authored pedagogical materials (private; de-identified). We do not publish corpus composition or counts.

Citation

@software{qlm_nto_classifier_2026, author = {Quantum Learning Machines}, title = {qlm-nto-classifier: pedagogical talk-move classification}, year = {2026}, url = {https://huggingface.co/QuantumLearningMachines/qlm-nto-classifier}, license = {Apache-2.0} }

Contact: hello@quantumlearningmachines.com · quantumlearningmachines.com

Downloads last month
17
Safetensors
Model size
0.2B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for QuantumLearningMachines/qlm-nto-classifier

Finetuned
(662)
this model