AI & ML interests
Text-to-SQL, Low-Resource NLP, Small Language Models (SLMs), Code-Mixed NLP, Cameroon Pidgin English, Camfranglais, Semantic Parsing
Recent Activity
Organization Card
Cameroon AI Research Lab (C-NLP)
Welcome to the official repository hub of the Cameroon AI Research Lab. We are an academic and community-driven research group dedicated to advancing Natural Language Processing (NLP), speech technology, and Artificial Intelligence resources for Cameroonian languages, West African Pidgin, and regional dialects.
🎯 Our Mission
- Language Preservation: Documenting and digitizing indigenous Cameroonian languages.
- Dataset Creation: Building high-quality, open-source benchmarks for low-resource languages.
- Model Adaptation: Fine-tuning small language models (SLMs) to make state-of-the-art AI efficient and accessible locally.
- Community Empowerment: Creating accessible AI tools to bridge the digital language divide across Central and West Africa.
📦 Key Projects
- Camspoken NLI: The first foundational Natural Language Inference dataset optimized for Cameroonic languages and Pidgin variants.
- Localized Text-to-SQL: Building benchmarks and models that translate Cameroonian languages and Pidgin queries into executable database scripts (SQL).
⚖️ Open Source Licensing
To maximize global collaboration, scientific progress, and commercial accessibility, the Cameroon AI Research Lab operates under an open-source framework:
- Unless stated otherwise, all software code, scripts, model architectures, and core assets published by this organization are licensed under the Apache License 2.0.
- Individual dataset file repositories may inherit this permissive Apache infrastructure or utilize data-specific permissive open licenses (like CC-BY-4.0) to ensure unrestricted community integration.
🤝 Contact & Collaborate
We welcome collaborations from computational linguists, machine learning engineers, students, and language enthusiasts.
- Email: tiani@tianipekins.com
- Affiliation: University of Buea, Cameroon
models 0
None public yet
datasets 0
None public yet