SLM-30M-QA
A fine-tuned version of slm-30m-base — a 30M-parameter decoder-only GPT built from scratch — adapted for question-answering. Built purely as a personal learning project.
Base Model
See slm-30m-base for architecture and pretraining details.
Fine-Tuning
- Data: Databricks Dolly-15k (filtered to short direct categories), Alpaca-Cleaned (~4,000 sampled examples), and Wikipedia-grounded Q&A pairs extracted from lead sentences
- Objective fix: corrected a target-alignment bug where tied embedding/output weights let the model learn a trivial "copy current token" shortcut; implemented proper next-token target shifting with
-100prompt masking - Checkpoint selection: instead of relying only on teacher-forced validation loss (which hides degeneration), checkpoints were evaluated with free-running generation on probe questions and only accepted if outputs were non-repetitive and non-blank
Usage
import torch
model = torch.load("finetuned.pt", map_location="cpu")
# prompts are wrapped as:
# ### Question:\n<your question>\n\n### Answer:\n
Inference uses temperature scaling, top-k sampling, and a repetition penalty to reduce degenerate/looping outputs.
Purpose
Built to understand instruction fine-tuning end-to-end — dataset mixing, loss masking, and evaluation beyond just validation loss — not intended as a production QA system.
Limitations
Small parameter count and limited fine-tuning data mean answers can be inconsistent or shallow compared to larger instruction-tuned models. Best treated as a learning artifact, not a usable assistant.