NIA LLM 110M - SFT (Instruction Fine‑tuned)
This is the instruction‑fine‑tuned version of the NIA 110M base model, built from scratch in Malawi 🇲🇼. This model is for research and experiment model
Model Details
Base Model:
- 110M parameters (10 layers, 768 dim, 12 heads)
- Pretrained on ~3.2B tokens
- Best pretrain loss: 3.3976
Fine‑tuning:
- Dataset: Mixed SQuAD (Q&A) + Alpaca (instructions)
- Steps: 4,110
- Final loss: 1.5845
- Learning rate: 3e-5
Capabilities
- Follows user instructions (chat format)
- Answers some factual questions (improved accuracy)
- Temperature sweet spot: 0.65
Usage
from transformers import AutoModelForCausalLM, AutoTokenizer
model = AutoModelForCausalLM.from_pretrained("Mathematicaljuice/nia")
tokenizer = AutoTokenizer.from_pretrained("Mathematicaljuice/nia")
prompt = "Explain photosynthesis in one paragraph."
inputs = tokenizer(prompt, return_tensors="pt")
output = model.generate(**inputs, max_new_tokens=100, temperature=0.65)
print(tokenizer.decode(output[0]))
Limitations
- Still hallucinates facts (especially uncommon ones)
- Cannot do arithmetic or math reasoning
- English only (for now)
- Small model size limits knowledge retention
- This model is for reseach only
Credits
Built by the Malawian ai start up based in blantyre Built by Lance Muyawa
License
Apache 2.0
- Downloads last month
- 118
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support