EDAN20 Assignment 4: CLD3-inspired language detector
This repository contains two small neural language classifiers trained on a
balanced subset of Tatoeba. The supported labels are Arabic (ara), French
(fra), Mandarin Chinese (cmn), Japanese (jpn), Korean (kor), English
(eng), Swedish (swe), and Danish (dan).
Both classifiers use relative frequencies of hashed, lowercased character unigrams, bigrams, and trigrams. The 2,583-dimensional input is passed through a 50-unit ReLU hidden layer and an eight-class output layer. This is inspired by CLD3, but it uses the frequency vector directly instead of learning a dense embedding for every hashed n-gram ID.
Files
nld.joblib: trained scikit-learnMLPClassifiernld.pth: PyTorchstate_dictfor a 2583--50--8 networknld_vectorizer.joblib: fittedDictVectorizernld_lang_codes.joblib: mapping from class indices to language codespredict.py: preprocessing and example inference for both models
Validation results
The deterministic split used 24,000 sentences (3,000 per language), with 80% for training and 20% for validation.
| Implementation | Accuracy | Macro F1 |
|---|---|---|
| scikit-learn | 0.9923 | 0.9922 |
| PyTorch | 0.9915 | 0.9914 |
The main remaining errors are Swedish--Danish confusions. Results should not be interpreted as performance on arbitrary web text: Tatoeba sentences are short, the eight classes are balanced, and no unseen languages are represented.
Example
python predict.py "Bonjour tout le monde"
The artifacts were produced with scikit-learn 1.0.2, PyTorch 2.5.0 (CPU), and NumPy 1.26.4.