NYEHASHAI

non-profit
Activity Feed

AI & ML interests

Rundi Machine Learning development and digital preservation of the kirundi language

Recent Activity

4Klanย  updated a Space 2 days ago
Nyehashai/README
4Klanย  updated a model 3 days ago
Nyehashai/Ntahokaja-0.31-Base
4Klanย  published a model 3 days ago
Nyehashai/Ntahokaja-0.31-Base
View all activity

Organization Card

Nyehash AI

Kirundi, kept legible to machines.

Nyehash AI builds datasets, models, and morphology-aware tools for Kirundi โ€” an agglutinative, tonal Bantu language spoken across Burundi and beyond.

Ishirahamwe ryo gukingira Ikirundi mu mashini nyabwonko


Why this exists

Most NLP tooling is built for isolating, non-agglutinative languages. Kirundi packs a noun's class, a verb's subject, tense, object, and voice into single inflected words โ€” a plain subword tokenizer breaks that structure apart at meaningless boundaries instead of meaningful ones.

We build morphology-aware tools, open datasets, and models that respect that structure, so Kirundi text stays usable โ€” for translation, search, education, and archival work โ€” instead of being flattened into noise.

A quick example

Word Segmentation Meaning
umuntu umu- (class 1) + -ntu person
abantu aba- (class 2) + -ntu people
barakina ba- (subj) + -ra- (tense) + -kin- (root) + -a they play
igitabo igi- (class 7) + -tabo book

What we build

  • Corpora โ€” collecting and cleaning Kirundi text (oral literature, news, dictionaries) into open, citable datasets.
  • Morphology tools โ€” rule-based analyzers and tokenizers built on Kirundi's noun-class and concord system, not borrowed defaults.
  • Language models โ€” fine-tunes and small models trained with Kirundi's grammar in mind, for translation and generation.

๐Ÿ“ฆ Datasets ยท ๐Ÿค– Models ยท โœ‰๏ธ nyehashi@gmail.com

ยฉ Nyehash AI โ€” digital preservation of the Kirundi language

datasets 0

None public yet