Tiny Vision Foundry
Tiny Vision Foundry compares a compact convolutional teacher with two copies of the same tiny student architecture:
- a student trained from labels alone;
- a student trained from labels plus the teacher's softened predictions.
The experiment uses the 8x8 handwritten-digit corpus bundled with scikit-learn. The data recipe converts it into deterministic train, validation, and test Parquet splits that load directly through Hugging Face Datasets.
Reproduce
uv run python projects/tiny-vision-foundry/prepare_data.py
uv run python projects/tiny-vision-foundry/train.py
All three networks use the same split and preprocessing. Final test metrics are computed only after training and model selection are complete.
Verified results
| Variant | Parameters | Test accuracy | Macro F1 |
|---|---|---|---|
| Teacher CNN | 13,706 | 97.78% | 0.9771 |
| Tiny student, labels only | 2,198 | 98.52% | 0.9848 |
| Tiny student, distilled | 2,198 | 97.41% | 0.9730 |
The tiny student is 6.24 times smaller than the teacher. In this first experiment, knowledge distillation reduced accuracy by 1.11 percentage points versus the identical student trained from labels alone. That result is preserved as a negative ablation; the labels-only student is the recommended artifact.