EDAN20 Assignment 4: CLD3-inspired language detector

This repository contains two small neural language classifiers trained on a balanced subset of Tatoeba. The supported labels are Arabic (ara), French (fra), Mandarin Chinese (cmn), Japanese (jpn), Korean (kor), English (eng), Swedish (swe), and Danish (dan).

Both classifiers use relative frequencies of hashed, lowercased character unigrams, bigrams, and trigrams. The 2,583-dimensional input is passed through a 50-unit ReLU hidden layer and an eight-class output layer. This is inspired by CLD3, but it uses the frequency vector directly instead of learning a dense embedding for every hashed n-gram ID.

Files

  • nld.joblib: trained scikit-learn MLPClassifier
  • nld.pth: PyTorch state_dict for a 2583--50--8 network
  • nld_vectorizer.joblib: fitted DictVectorizer
  • nld_lang_codes.joblib: mapping from class indices to language codes
  • predict.py: preprocessing and example inference for both models

Validation results

The deterministic split used 24,000 sentences (3,000 per language), with 80% for training and 20% for validation.

Implementation Accuracy Macro F1
scikit-learn 0.9923 0.9922
PyTorch 0.9915 0.9914

The main remaining errors are Swedish--Danish confusions. Results should not be interpreted as performance on arbitrary web text: Tatoeba sentences are short, the eight classes are balanced, and no unseen languages are represented.

Example

python predict.py "Bonjour tout le monde"

The artifacts were produced with scikit-learn 1.0.2, PyTorch 2.5.0 (CPU), and NumPy 1.26.4.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support