MicroT-test1-50K-SmolTalk2
|
Micro Transformer — Reference Baseline Vanilla Attention • RoPE • Byte-Level • Decoder-Only |
📋 Overview
MicroT-test1-50K-SmolTalk2 is a 49,888-parameter vanilla decoder-only transformer — multi-head causal self-attention with RoPE — pretrained on SmolTalk2 conversation data (instead of the project's Discord-Dialogues baseline) and then fine-tuned with FMSP on 9,012 general-knowledge QA pairs.
This is the 50K member of the MicroT-test1 family: the registered attention-based reference baseline of the MicroMixer-4 project, here rerun on SmolTalk2 as part of the dataset-efficiency comparison study — six parameter budgets × two architectures × two open pretraining corpora, all fine-tuned with the identical P05 FMSP recipe at seed 42. Analysis.
It is deliberately boring — the standard 2018–2020 transformer recipe parameterized down to the sub-1M regime: no flash attention, no SwiGLU, no ALiBi, no QKNorm, no sliding window, no MQA/GQA, no MoE.
🏗️ Architecture
graph TD
A[Byte Input] --> B[Embed 256→32]
B --> C[Transformer Block × 3]
C --> D[RMSNorm]
D --> E[LM Head Tied with Embed]
E --> F[Byte Output]
subgraph "Transformer Block (pre-norm)"
X[Input 32] --> N1[RMSNorm]
N1 --> AT["MHA 2 heads × d_head 16<br/>RoPE θ=10000 on q,k · causal SDPA"]
AT --> R1[+ residual]
R1 --> N2[RMSNorm]
N2 --> MLP["GELU MLP 32→152→32"]
MLP --> R2[+ residual]
end
style A fill:#007BFF,color:#fff
style F fill:#00D620,color:#fff
style AT fill:#FF6600,color:#fff
Model Configuration
| Parameter | Value |
|---|---|
| Hidden Dimension (d_model) | 32 |
| Attention Heads | 2 (d_head = 16 at every size) |
| Number of Blocks | 3 |
| FFN Hidden | 152 |
| Position Encoding | RoPE θ=10000 on q/k only (non-persistent buffers) |
| Attention | Causal MHA via F.scaled_dot_product_attention(is_causal=True) |
| Activation | GELU |
| Biases | None — no bias parameters anywhere |
| Normalization | RMSNorm (pre-norm) |
| Max Sequence Length | 1024 |
| Vocabulary Size | 256 (byte-level) |
| Output Head | Tied with input embedding |
Core Components
┌──────────────────────────────────────────────┐
│ Transformer Block (×3) │
│ h = h + MHA(RMSNorm(h)) # RoPE q/k, causal│
│ h = h + MLP(RMSNorm(h)) # GELU d→ffn→d │
│ no biases, no flash, no tricks — vanilla │
└──────────────────────────────────────────────┘
The d_head=16 contract is hard-asserted across all six sizes so that attention-head behavior is comparable at every budget and never confounds the memorization measurements.
🎯 Generation Examples
Questions the model was trained on (FMSP train set, 9,012 QA pairs — greedy decoding, repetition_penalty=1.2, no_repeat_ngram_size=4):
[Prompt] User: Who painted the Mona Lisa?
Assistant:
[Output] The Sarricate against of the American internet in the Anderson of the Africa of the Africa of t…
fabricates — gives a wrong answer to a trained question at this size
[Prompt] User: Who painted The Starry Night?
Assistant:
[Output] The Sarricate is a series of the consider that contains of the consider than the largest contai…
fabricates — gives a wrong answer to a trained question at this size
Questions the model has never seen and cannot answer (unanswerable probe — the correct behavior is to decline; the model's actual behavior is shown):
[Prompt] User: Who painted the Glimmering Frostberry?
Assistant:
[Output] The Sarricate is a series of the consider that contains of the consider than the largest contai…
fabricates — plausible-sounding nonsense on a nonexistent subject
[Prompt] User: Who composed the Symphony of Hollow Dawn?
Assistant:
[Output] The Sarricate is a series of the consider that contains of the consider than the largest contai…
fabricates — plausible-sounding nonsense on a nonexistent subject
📊 Results
Pretraining (SmolTalk2, V76 recipe, 3 epochs)
| Metric | 1 ep | 2 ep | 3 ep |
|---|---|---|---|
| Val PPL | 4.04 | 3.91 | 3.68 |
Identical recipe to the MicroMixer-4 mixer: AdamW lr 3e-3 · WSD (warmup 500) · wd 0.01 · bs 16 · seq 1024 · seed 42.
FMSP fine-tuning (small-qa-en-10k, P05 recipe, 10 epochs)
| Metric | Value |
|---|---|
| Train QA pairs | 9,012 |
| Held-out QA pairs | 988 |
| Best-val checkpoint | fmsp_epoch_9.safetensors (val loss 1.1292) |
| freeze_fraction | 0.05 (true freeze) |
| Loss | answer-only CE + probe KL (weight 0.5) |
Evaluation battery (post-FMSP)
| Axis | MicroT-test1-50K-SmolTalk2 |
|---|---|
| Chatter fluency d2 (cycles) | 0.549 (6/9) |
| Full-988 EM (seed 42) | 0 |
| Q-relevance echo / hijack % | 8.0 / 49.0 |
| OOD hijack % | 50.8% |
| Unanswerable fabrication /18 | 14 |
| Discord PPL | 10.20 |
Single-seed run (seed 42); the discord-pretrained cards report a 3-seed mean for Full-988 EM. ‡ where marked: degenerate-pass.
MicroT-test1 SmolTalk2 family (same protocol, all sizes)
| Size | Params | 3ep Val PPL | Chatter d2 | Full-988 EM | qrel echo/hijack | OOD hijack |
|---|---|---|---|---|---|---|
| 1M | 996,736 | 2.12 | 0.911 | 744 | 100.0 / 0.0 | 37.3% |
| 500K | 498,528 | 2.32 | 0.932 | 580 | 86.0 / 7.0 | 42.4% |
| 300K | 297,680 | 2.53 | 0.775 | 137 | 55.0 / 26.0 | 40.7% |
| 100K | 97,872 | 3.13 | 0.673 | 0 | 5.0 / 86.0 | 55.9% |
| 50K | 49,888 | 3.68 | 0.549 | 0 | 8.0 / 49.0 | 50.8% |
| 10K | 9,808 | 6.04 | 0.447 | 0 | 0.0 / 0.0 | 0.0% |
Seed-42 single runs (pretrained on SmolTalk2; the discord-pretrained families report 3-seed EM means).
📚 Training Data
- Pretraining: SmolTalk2 — 200K-cap sample of the
smoltalk_smollm3_smol_magpie_ultra_no_thinksplit (200K of 406,843; Apache 2.0), ~6-turn multi-turn conversations with no reasoning traces and system messages excluded, flattened toUser:/Assistant:format, 1024-byte sequences, 3 epochs. - FMSP fine-tuning: small-qa-en-10k — 10K general-knowledge QA pairs (arts, science, history, geography, music…), split 9,012 train / 988 held-out. 10 epochs under the P05 recipe (5% of parameters frozen-true, answer-only CE, probe-KL 0.5).
🔧 Usage
Files in this repository
fmsp_epoch_{0..9}.safetensors— per-epoch FMSP weights (pickle-free safetensors).fmsp_epoch_9.safetensorsis the best-val checkpoint for this size.
Load and generate (local clone)
import torch
from safetensors.torch import load_file
from src.model_v88_transformer import MicroMixerV88Transformer, v88_transformer_50k
from src.fmsp import attach_adapter
from src.tokenizer import ByteTokenizer
# Clone the code repository first:
# git clone https://github.com/llaa33219/MicroMixer-4.git && cd MicroMixer-4
cfg = v88_transformer_50k()
model = MicroMixerV88Transformer(cfg)
attach_adapter(model, d_model=cfg.d_model, rank=16) # FMSP adapter (trained weights are in the file)
model.load_state_dict(load_file("fmsp_epoch_9.safetensors"), strict=True)
model.eval()
tok = ByteTokenizer()
prompt = "User: Who painted the Mona Lisa?\n\nAssistant: "
ids = tok.encode(prompt)
if ids and ids[-1] == tok.eos_token_id:
ids = ids[:-1] # ByteTokenizer appends EOS; the prompt must end open
ids = torch.tensor([ids])
with torch.no_grad():
out = model.generate(
ids, max_new_tokens=200,
temperature=0.0, # greedy — used for all reported numbers
repetition_penalty=1.2,
no_repeat_ngram_size=4,
eos_token_id=tok.eos_token_id,
)
print(tok.decode(out[0].tolist()))
Load from Hugging Face Hub (no clone of the weights needed)
import torch
from huggingface_hub import hf_hub_download
from safetensors.torch import load_file
from src.model_v88_transformer import MicroMixerV88Transformer, v88_transformer_50k
from src.fmsp import attach_adapter
REPO = "llaa33219/MicroT-test1-50K-SmolTalk2"
cfg = v88_transformer_50k()
model = MicroMixerV88Transformer(cfg)
attach_adapter(model, d_model=cfg.d_model, rank=16)
model.load_state_dict(
load_file(hf_hub_download(REPO, "fmsp_epoch_9.safetensors")), strict=True)
model.eval()
# ... generate as above
⚠️ Limitations
| Limitation | Description |
|---|---|
| Micro parameters | 49,888 parameters; capacity is the binding constraint on every axis |
| Reference baseline, not a product | Exists to score the Mixer against attention at matched budget |
| Knows only what it memorized | Knowledge is limited to the 9,012 trained QA pairs + the pretraining-corpus distribution |
| Does not abstain | Unknown questions are answered with fabrication or degeneration, not refusal — see the examples above |
| Byte-level noise | 256-vocab byte tokenizer; PPL not comparable to BPE baselines |
| Research use only | Architecture/scaling research artifact, not a production model |
🧬 Context
This is the 50K SmolTalk2-pretrained arm of the dataset-efficiency comparison study (24 arms: {V87, V88} × {UltraChat, SmolTalk2} × six sizes) in the MicroMixer-4 project. Each arm swaps the Discord-Dialogues pretraining baseline for an open corpus (SmolTalk2) and reruns the identical P05 FMSP recipe at seed 42. Sibling repos: llaa33219/MicroMixer-4-{1M..10K}-{UltraChat,SmolTalk2} and llaa33219/MicroT-test1-{1M..10K}-{UltraChat,SmolTalk2}; the discord-pretrained baselines are llaa33219/MicroMixer-4-{1M..10K} and llaa33219/MicroT-test1-{1M..10K}. Full analysis: DATASET_COMPARISON_ANALYSIS.md.
Part of the MicroMixer-4 research project — V88 transformer reference (MicroT-test1), 50K preset, SmolTalk2 pretraining