SLM-30M-QA

A fine-tuned version of slm-30m-base — a 30M-parameter decoder-only GPT built from scratch — adapted for question-answering. Built purely as a personal learning project.

Base Model

See slm-30m-base for architecture and pretraining details.

Fine-Tuning

  • Data: Databricks Dolly-15k (filtered to short direct categories), Alpaca-Cleaned (~4,000 sampled examples), and Wikipedia-grounded Q&A pairs extracted from lead sentences
  • Objective fix: corrected a target-alignment bug where tied embedding/output weights let the model learn a trivial "copy current token" shortcut; implemented proper next-token target shifting with -100 prompt masking
  • Checkpoint selection: instead of relying only on teacher-forced validation loss (which hides degeneration), checkpoints were evaluated with free-running generation on probe questions and only accepted if outputs were non-repetitive and non-blank

Usage

import torch
model = torch.load("finetuned.pt", map_location="cpu")
# prompts are wrapped as:
# ### Question:\n<your question>\n\n### Answer:\n

Inference uses temperature scaling, top-k sampling, and a repetition penalty to reduce degenerate/looping outputs.

Purpose

Built to understand instruction fine-tuning end-to-end — dataset mixing, loss masking, and evaluation beyond just validation loss — not intended as a production QA system.

Limitations

Small parameter count and limited fine-tuning data mean answers can be inconsistent or shallow compared to larger instruction-tuned models. Best treated as a learning artifact, not a usable assistant.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support