Exploring Pedagogical Alignment of LLMs Using Student Errors and Tutor Moves โ€” trained models

All models trained for the thesis Exploring Pedagogical Alignment of LLMs Using Student Errors and Tutor Moves (Jaeyeop Chung, Technical University of Munich, 2026). Code, data and notebooks: https://github.com/JaeyeopC/Thesis_Repository. Each subfolder is a complete model with its own model card.

Subfolder Model Base Role in the thesis
llama3-8b-instruct-sft-adapter/ LoRA adapter, SFT on 8,111 human-authored MathDial tutor turns Llama-3-8B-Instruct starting point and reference policy for DPO
llama3-8b-instruct-dpo-adapter-top10/ LoRA adapter, DPO on the top 10 % preference pairs per error type (1,940 training pairs) Llama-3-8B-Instruct DPO 10 %
llama3-8b-instruct-dpo-adapter-top30/ LoRA adapter, DPO on the top 30 % (5,828 training pairs) Llama-3-8B-Instruct DPO 30 %
llama3-8b-instruct-dpo-adapter-top50/ LoRA adapter, DPO on the top 50 % (9,719 training pairs) Llama-3-8B-Instruct DPO 50 %
roberta-base-tutor-move-classifier/ Tutor-move classifier (focus / generic / probing / telling), tutor-turn-only input, AUC maximization RoBERTa-base tutor-move posteriors for scoring candidate responses
tutor-move-classifiers/ (optional) All 18 classifier configurations (encoder ร— input ร— objective) RoBERTa-base classifier comparison, Section 5.1

How to use

The Llama-3 base model is gated: request access to meta-llama/Meta-Llama-3-8B-Instruct on the Hub and log in with hf auth login.

from transformers import AutoModelForCausalLM, AutoTokenizer, pipeline
from peft import PeftModel

repo = "woduq132/thesis_repository"

# tutor model: base + one of the adapters
sub = "llama3-8b-instruct-dpo-adapter-top30"
tokenizer = AutoTokenizer.from_pretrained(repo, subfolder=sub)
model = AutoModelForCausalLM.from_pretrained("meta-llama/Meta-Llama-3-8B-Instruct", dtype="bfloat16", device_map="auto")
model = PeftModel.from_pretrained(model, repo, subfolder=sub)

# tutor-move classifier
from transformers import AutoModelForSequenceClassification
cls_sub = "roberta-base-tutor-move-classifier"
clf = pipeline("text-classification",
               model=AutoModelForSequenceClassification.from_pretrained(repo, subfolder=cls_sub),
               tokenizer=AutoTokenizer.from_pretrained(repo, subfolder=cls_sub), top_k=None)

Evaluation summary

Desired Annotation Match Rate (DAMR, %) of the five tutor-model variants on the 298 test contexts, judged by gpt-4.1-mini (share of responses whose label equals the desired label of each pedagogical dimension; see Section 5.3 of the thesis):

Dimension (desired label) Base SFT DPO 10 % DPO 30 % DPO 50 %
Mistake identification (Yes) 95.3 41.6 63.1 65.1 83.9
Mistake location (Yes) 93.3 38.9 50.7 51.3 68.8
Revealing the answer (No) 98.7 94.0 70.1 88.6 80.9
Providing guidance (Yes) 88.6 28.2 14.4 22.8 27.2
Actionability (Yes) 88.6 29.9 14.4 27.2 28.9
Coherence (Yes) 99.7 97.7 90.9 95.3 96.3
Tutor tone (Encouraging) 100.0 69.8 24.8 49.7 34.9
Human-likeness (Yes) 99.7 84.6 48.3 72.8 62.4

Selected tutor-move classifier, test split: accuracy 0.736, macro-F1 0.642, macro ROC-AUC 0.864, MCC 0.616.

Licenses

The Llama-3 adapters are released under the Llama 3 Community License of the base model (Built with Meta Llama 3); the RoBERTa-base classifiers under the MIT license.

Citation

Exploring Pedagogical Alignment of LLMs Using Student Errors and Tutor Moves โ€” Jaeyeop Chung, M.Sc. thesis, Technical University of Munich, 2026. Code and data: https://github.com/JaeyeopC/Thesis_Repository

Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for woduq132/thesis_repository

Adapter
(326)
this model