NYEHASHAI
AI & ML interests
Rundi Machine Learning development and digital preservation of the kirundi language
Recent Activity
Nyehash AI
Kirundi, kept legible to machines.
Nyehash AI builds datasets, models, and morphology-aware tools for Kirundi โ an agglutinative, tonal Bantu language spoken across Burundi and beyond.
Ishirahamwe ryo gukingira Ikirundi mu mashini nyabwonko
Why this exists
Most NLP tooling is built for isolating, non-agglutinative languages. Kirundi packs a noun's class, a verb's subject, tense, object, and voice into single inflected words โ a plain subword tokenizer breaks that structure apart at meaningless boundaries instead of meaningful ones.
We build morphology-aware tools, open datasets, and models that respect that structure, so Kirundi text stays usable โ for translation, search, education, and archival work โ instead of being flattened into noise.
A quick example
| Word | Segmentation | Meaning |
|---|---|---|
umuntu |
umu- (class 1) + -ntu |
person |
abantu |
aba- (class 2) + -ntu |
people |
barakina |
ba- (subj) + -ra- (tense) + -kin- (root) + -a |
they play |
igitabo |
igi- (class 7) + -tabo |
book |
What we build
- Corpora โ collecting and cleaning Kirundi text (oral literature, news, dictionaries) into open, citable datasets.
- Morphology tools โ rule-based analyzers and tokenizers built on Kirundi's noun-class and concord system, not borrowed defaults.
- Language models โ fine-tunes and small models trained with Kirundi's grammar in mind, for translation and generation.
๐ฆ Datasets ยท ๐ค Models ยท โ๏ธ nyehashi@gmail.com
ยฉ Nyehash AI โ digital preservation of the Kirundi language