NIA LLM 110M - SFT (Instruction Fine‑tuned)

This is the instruction‑fine‑tuned version of the NIA 110M base model, built from scratch in Malawi 🇲🇼. This model is for research and experiment model

Model Details

Base Model:

  • 110M parameters (10 layers, 768 dim, 12 heads)
  • Pretrained on ~3.2B tokens
  • Best pretrain loss: 3.3976

Fine‑tuning:

  • Dataset: Mixed SQuAD (Q&A) + Alpaca (instructions)
  • Steps: 4,110
  • Final loss: 1.5845
  • Learning rate: 3e-5

Capabilities

  • Follows user instructions (chat format)
  • Answers some factual questions (improved accuracy)
  • Temperature sweet spot: 0.65

Usage

from transformers import AutoModelForCausalLM, AutoTokenizer

model = AutoModelForCausalLM.from_pretrained("Mathematicaljuice/nia")
tokenizer = AutoTokenizer.from_pretrained("Mathematicaljuice/nia")

prompt = "Explain photosynthesis in one paragraph."
inputs = tokenizer(prompt, return_tensors="pt")
output = model.generate(**inputs, max_new_tokens=100, temperature=0.65)
print(tokenizer.decode(output[0]))

Limitations

  • Still hallucinates facts (especially uncommon ones)
  • Cannot do arithmetic or math reasoning
  • English only (for now)
  • Small model size limits knowledge retention
  • This model is for reseach only

Credits

Built by the Malawian ai start up based in blantyre Built by Lance Muyawa

License

Apache 2.0

Downloads last month
118
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support