tohio/slm-125m-base

tohio/slm-125m-base is a small language model (125M parameters) built from scratch and aligned end-to-end using the slm-gpt engine.

Architecture Highlights

  • Decoder-Only Transformer with Rotary Position Embeddings (RoPE)
  • Attention: Grouped-Query Attention (GQA, 12:4 ratio)
  • Activation: SwiGLU Feed-Forward Network ($d_{ffn} = 2048$)
  • Normalization: Bias-free Pre-LayerNorm / RMSNorm
  • Weight Tying: Tied input embedding (embed_tokens) and output head (lm_head)
  • Hugging Face Native: Directly compatible with LlamaForCausalLM

Pre-training Curriculum

The base model was pre-trained across an interleaved multi-source domain mixture composed of:

  • FineWeb-Edu
  • DCLM-Edu
  • The Stack-Edu
  • NuminaMath-CoT
  • OpenMathReasoning
  • SLM-Synthetic-Pretrain

Alignment Lineage

  1. Pre-training: Multi-GPU Distributed Data Parallel (DDP) over rank-disjoint memory-mapped token shards.
  2. Supervised Fine-Tuning (SFT): Full parameter instruction-tuning on ChatML formatted dialogues with prompt loss masking (ignore_index=-100).
  3. Direct Preference Optimization (DPO): Single-stage pairwise preference alignment optimizing chosen vs. rejected generations.

Prompt Format (ChatML)

This model adheres strictly to the ChatML template:

<|im_start|>system
You are a helpful AI assistant.<|im_end|>
<|im_start|>user
Write a Python script to compute the Fibonacci sequence efficiently.<|im_end|>
<|im_start|>assistant

Quick Start via transformers

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

model_id = "tohio/slm-125m-base"

tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
    model_id,
    torch_dtype=torch.bfloat16,
    device_map="auto",
)

messages = [
    {"role": "system", "content": "You are a concise AI assistant."},
    {"role": "user", "content": "Explain quantum computing in one sentence."},
]

prompt = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
inputs = tokenizer(prompt, return_tensors="pt").to(model.device)

outputs = model.generate(**inputs, max_new_tokens=100, temperature=0.7, top_p=0.9)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))

Trained and published autonomously via slm-gpt.

Downloads last month
124
Safetensors
Model size
0.2B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Datasets used to train tohio/slm-125m-base-legacy