leap-kt · IEKT
IEKT —
Part of leap-kt-toolkit, a systematic re-implementation of published Knowledge Tracing models under one protocol. This repository holds every fold of every dataset this model has been run on, with the per-epoch training logs and the exact user split alongside the checkpoints.
Protocol
User-level 80/20 train/test split · 5-fold cross-validation over the training portion · held-out fold as validation · early stopping patience 10 on validation AUC · max 200 epochs.
Every cell in the project runs under identical settings; a cell that cannot is recorded as a documented failure rather than re-run under bespoke settings.
Results
| dataset | AUC | ACC | F1 | published reference | delta |
|---|---|---|---|---|---|
algebra2005 |
0.8148 ± 0.0014 | 0.8110 | 0.8809 | — | — |
assist2009 |
0.7714 ± 0.0025 | 0.7383 | 0.8116 | — | — |
assist2012 |
0.7608 ± 0.0022 | 0.7482 | 0.8322 | — | — |
assist2015 |
0.7297 ± 0.0011 | 0.7524 | 0.8480 | — | — |
bridge2algebra2006 |
0.7858 ± 0.0022 | 0.8444 | 0.9126 | — | — |
dbe_kt22 |
0.8068 ± 0.0024 | 0.7972 | 0.8756 | — | — |
ednet500 |
0.7051 ± 0.0022 | 0.6863 | 0.7749 | — | — |
Per-fold values are in each dataset's summary.json. The mean is never reported without the spread — 0.75 ± 0.001 and 0.75 ± 0.09 are different claims.
Why these numbers may differ from other reproductions
Multi-concept questions are not expanded into multiple rows. Toolkits that do expand them place consecutive test positions carrying the same question and the same response, so a model is shown the answer one step before predicting it; on ASSIST2009 that is around 37% of positions and lifts DKT from a published ~0.75 to ~0.89 AUC. Here concepts are an extra axis on the interaction rather than extra rows, so the leak is not expressible and every interaction is scored exactly once.
Cells carrying a published reference value are additionally leak-audited before release: train/test user disjointness, no window crossing the split boundary, exactly-once scoring, and a label-shuffle control that must collapse AUC to chance. Cells with no comparable published number rely on the structural guarantee above rather than on that audit.
Files
<dataset>/summary.json mean ± std and per-fold AUC
<dataset>/split.json the exact user partition, with a checksum
<dataset>/fold<k>/checkpoint/ config.json + weights
<dataset>/fold<k>/epochs.jsonl every epoch's train loss and validation metrics;
each row carries its own model/dataset/fold
<dataset>/fold<k>/run.json protocol and package version for that run
Provenance
Produced by leap-kt at commit(s) 2136c89, 40bb469, bf03379.