MicroMixer-4 Logo

MicroMixer-4-100K-SmolTalk2

Parameters Architecture FMSP

Micro Language Model
Attention-Free • MLP-Only • Byte-Level • Content-Gated Dilated Convolution

GitHub

📋 Overview

MicroMixer-4-100K-SmolTalk2 is a 95,084-parameter pure MLP-Mixer causal language model — no attention, no recurrence, no SSM — pretrained on SmolTalk2 conversation data (instead of the project's Discord-Dialogues baseline) and then fine-tuned with FMSP on 9,012 general-knowledge QA pairs.

This is the 100K member of the MicroMixer-4 (V87 Final) family: part of the project's dataset-efficiency comparison study — six parameter budgets × two architectures × two open pretraining corpora (UltraChat 200k and SmolTalk2), all fine-tuned with the identical P05 FMSP recipe at seed 42. Analysis.

The backbone is V87 Final, the project's champion architecture — a CCD-Mixer (Content-gated mixture of shared-weight Dilated convolutions) crowned overall champion of the 1M architecture census (V86), frozen as the final chassis and scaled to six parameter budgets. The 100K preset reproduces the champion recipe verbatim at its budget.

🏗️ Architecture

graph TD
    A[Byte Input] --> B[Embed 256→48 NoPE]
    B --> C[CCD-Mixer Block × 4]
    C --> D[RMSNorm]
    D --> E[LM Head Tied with Embed]
    E --> F[Byte Output]

    subgraph "CCD-Mixer Block"
        X[Input 48] --> U["Linear d→2d → split v, g"]
        U --> RP[Full RoPE on v AND g]
        RP --> M["Shared-weight dilated conv<br/>dilations 1·2·4·8, k=65"]
        M --> G["Per-position 4-way gate<br/>softmax(Linear_dil(x)/τ)"]
        G --> O["W_o(v ⊙ g)  — zero-init"]
        O --> SW[SwiGLU Channel-Mix]
        SW --> RM[ReMixerLayer sidecar]
    end

    style A fill:#007BFF,color:#fff
    style F fill:#00D620,color:#fff
    style G fill:#AE00FF,color:#fff
    style M fill:#FF6600,color:#fff

Model Configuration

Parameter Value
Hidden Dimension (d_model)48
Number of Blocks4
Token-MixGLCTokenMixCCD (content-gated mixture of shared-weight dilated causal conv)
Dilations(1, 2, 4, 8) — one shared depthwise kernel, zero extra conv params
Depthwise Kernel Size65
RoPEFull RoPE on both v and g (V76 "RPG" pattern)
Channel-MixSwiGLU
SidecarReMixerLayer per block (label_dim 16, pool_heads 4)
Max Sequence Length1024
Vocabulary Size256 (byte-level)
Position EncodingRoPE inside token-mix only; no position embedding table
NormalizationRMSNorm (pre-norm)
Output HeadTied with input embedding
Zero-InitW_o, dil_gate, log_τ — silent at init

Core Components

┌──────────────────────────────────────────────────────────────┐
│                 CCD-Mixer Block (×4)                        │
│  u    = Linear(d → 2d)(x)                                    │
│  v, g = u.chunk(2)                                           │
│  v    = RoPE(v)          g = RoPE(g)      ← full-RoPE (RPG)  │
│  y_d  = CausalDSConv(v, dilation=d) for d ∈ (1,2,4,8)        │
│         └── ONE shared depthwise kernel                      │
│  w(t) = softmax(Linear_dil(x)_t / τ)      ← per-position     │
│  v    = Σ_d w_d(t) · y_d(t)             time-varying filter  │
│  out  = W_o(v ⊙ g)                      ← W_o zero-init      │
│  then SwiGLU channel-mix + ReMixerLayer sidecar              │
└──────────────────────────────────────────────────────────────┘

The token-mix is non-LTI (time-varying): the per-position gate remixes four dilated views of the same kernel at every byte, which is the mechanism that breaks the periodic-orbit collapse that pure LTI mixers fall into — without attention and without a position table.


🎯 Generation Examples

Questions the model was trained on (FMSP train set, 9,012 QA pairs — greedy decoding, repetition_penalty=1.2, no_repeat_ngram_size=4):

[Prompt] User: Who painted the Mona Lisa?
Assistant:
[Output] The capital city of Giek I was the large-difference in the world was a red three painting the c…

fabricates — gives a wrong answer to a trained question at this size

[Prompt] User: Who painted The Starry Night?
Assistant:
[Output] The Main Dance was a Proteineas Brazil in 1971 and the Hand and Earth and the Happending Africa…

fabricates — gives a wrong answer to a trained question at this size

Questions the model has never seen and cannot answer (unanswerable probe — the correct behavior is to decline; the model's actual behavior is shown):

[Prompt] User: Who painted the Glimmering Frostberry?
Assistant:
[Output] The capital city of Greek in 1999 and 1900 to 1930 to 1930 to 1930.

fabricates — plausible-sounding nonsense on a nonexistent subject

[Prompt] User: Who composed the Symphony of Hollow Dawn?
Assistant:
[Output] The Sun is the country in the America was the famous Europe that is the most of the United Stat…

fabricates — plausible-sounding nonsense on a nonexistent subject


📊 Results

Pretraining (SmolTalk2, V76 recipe, 3 epochs)

Metric 1 ep 2 ep 3 ep
Val PPL 3.46 3.38 3.16

AdamW lr 3e-3 · WSD (warmup 500) · wd 0.01 · bs 16 · seq 1024 · seed 42 · plain CE on non-pad bytes.

FMSP fine-tuning (small-qa-en-10k, P05 recipe, 10 epochs)

Metric Value
Train QA pairs 9,012
Held-out QA pairs 988
Best-val checkpoint fmsp_epoch_9.safetensors (val loss 0.7490)
freeze_fraction 0.05 (true freeze)
Loss answer-only CE + probe KL (weight 0.5)

Evaluation battery (post-FMSP)

Axis MicroMixer-4-100K-SmolTalk2
Chatter fluency d2 (cycles) 0.518 (6/9)
Full-988 EM (seed 42) 0
Q-relevance echo / hijack % 7.0 / 76.0
OOD hijack % 66.1%
Unanswerable fabrication /18 17
Discord PPL 15.13

Single-seed run (seed 42); the discord-pretrained cards report a 3-seed mean for Full-988 EM. ‡ where marked: degenerate-pass.

MicroMixer-4 SmolTalk2 family (same protocol, all sizes)

Size Params 3ep Val PPL Chatter d2 Full-988 EM qrel echo/hijack OOD hijack
1M 996,873 2.23 0.960 739 100.0 / 0.0 47.5%
500K 491,742 2.42 0.906 481 61.0 / 32.0 49.2%
300K 292,525 2.61 0.900 133 18.0 / 63.0 57.6%
100K 95,084 3.16 0.518 0 7.0 / 76.0 66.1%
50K 48,684 3.66 0.480 0 9.0 / 49.0 64.4%
10K 9,666 6.00 0.426 0 0.0 / 0.0 0.0%

Seed-42 single runs (pretrained on SmolTalk2; the discord-pretrained families report 3-seed EM means).


📚 Training Data

  1. Pretraining: SmolTalk2 — 200K-cap sample of the smoltalk_smollm3_smol_magpie_ultra_no_think split (200K of 406,843; Apache 2.0), ~6-turn multi-turn conversations with no reasoning traces and system messages excluded, flattened to User:/Assistant: format, 1024-byte sequences, 3 epochs.
  2. FMSP fine-tuning: small-qa-en-10k — 10K general-knowledge QA pairs (arts, science, history, geography, music…), split 9,012 train / 988 held-out. 10 epochs under the P05 recipe (5% of parameters frozen-true, answer-only CE, probe-KL 0.5).

🔧 Usage

Files in this repository

  • fmsp_epoch_{0..9}.safetensors — per-epoch FMSP weights (pickle-free safetensors). fmsp_epoch_9.safetensors is the best-val checkpoint for this size.

Load and generate (local clone)

import torch
from safetensors.torch import load_file
from src.model_v87_final import MicroMixerV87Final, v87_final_100k
from src.fmsp import attach_adapter
from src.tokenizer import ByteTokenizer

# Clone the code repository first:
# git clone https://github.com/llaa33219/MicroMixer-4.git && cd MicroMixer-4

cfg = v87_final_100k()
model = MicroMixerV87Final(cfg)
attach_adapter(model, d_model=cfg.d_model, rank=16)   # FMSP adapter (trained weights are in the file)
model.load_state_dict(load_file("fmsp_epoch_9.safetensors"), strict=True)
model.eval()

tok = ByteTokenizer()
prompt = "User: Who painted the Mona Lisa?\n\nAssistant: "
ids = tok.encode(prompt)
if ids and ids[-1] == tok.eos_token_id:
    ids = ids[:-1]                       # ByteTokenizer appends EOS; the prompt must end open
ids = torch.tensor([ids])
with torch.no_grad():
    out = model.generate(
        ids, max_new_tokens=200,
        temperature=0.0,                 # greedy — used for all reported numbers
        repetition_penalty=1.2,
        no_repeat_ngram_size=4,
        eos_token_id=tok.eos_token_id,
    )
print(tok.decode(out[0].tolist()))

Load from Hugging Face Hub (no clone of the weights needed)

import torch
from huggingface_hub import hf_hub_download
from safetensors.torch import load_file
from src.model_v87_final import MicroMixerV87Final, v87_final_100k
from src.fmsp import attach_adapter

REPO = "llaa33219/MicroMixer-4-100K-SmolTalk2"

cfg = v87_final_100k()
model = MicroMixerV87Final(cfg)
attach_adapter(model, d_model=cfg.d_model, rank=16)
model.load_state_dict(
    load_file(hf_hub_download(REPO, "fmsp_epoch_9.safetensors")), strict=True)
model.eval()
# ... generate as above

⚠️ Limitations

Limitation Description
Micro parameters 95,084 parameters; capacity is the binding constraint on every axis
Knows only what it memorized Knowledge is limited to the 9,012 trained QA pairs + the pretraining-corpus distribution
Does not abstain Unknown questions are answered with fabrication or degeneration, not refusal — see the examples above
Byte-level noise 256-vocab byte tokenizer; PPL not comparable to BPE baselines
Research use only Architecture/scaling research artifact, not a production model

🧬 Context

This is the 100K SmolTalk2-pretrained arm of the dataset-efficiency comparison study (24 arms: {V87, V88} × {UltraChat, SmolTalk2} × six sizes) in the MicroMixer-4 project. Each arm swaps the Discord-Dialogues pretraining baseline for an open corpus (SmolTalk2) and reruns the identical P05 FMSP recipe at seed 42. Sibling repos: llaa33219/MicroMixer-4-{1M..10K}-{UltraChat,SmolTalk2} and llaa33219/MicroT-test1-{1M..10K}-{UltraChat,SmolTalk2}; the discord-pretrained baselines are llaa33219/MicroMixer-4-{1M..10K} and llaa33219/MicroT-test1-{1M..10K}. Full analysis: DATASET_COMPARISON_ANALYSIS.md.


GitHub

Part of the MicroMixer-4 research project — V87 Final (CCD-Mixer) family, 100K preset, SmolTalk2 pretraining

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Datasets used to train llaa33219/MicroMixer-4-100K-SmolTalk2

Collection including llaa33219/MicroMixer-4-100K-SmolTalk2