🦠 MicroSupra-10k

Recreation of SupraLabs/MicroSupra-1k with exactly 10,000 parameters, trained on a mix of Fineweb-Edu + DCLM-Edu, entirely on CPU.

Model Config

  • Parameters: 10,000 (exactly!)
  • Architecture: LLaMA (LlamaForCausalLM)
  • Vocab size: 1024 (BPE tokenizer reused from the original)
  • Hidden Size: 8
  • Intermediate Size: 69
  • Head Dim: 4 (1 head, 1 layer)
  • Max Position Embeddings: 256
  • Tied embeddings: true
  • Learning rate: 5e-3 (cosine decay β†’ 5e-4, warmup 200 steps)

Data

Mix 50/50 (interleaved), tokenized on the fly (streaming):

  • HuggingFaceFW/fineweb-edu β€” config sample-10BT
  • HuggingFaceTB/dclm-edu β€” filtered to edu_int_score >= 3 (as recommended for small models)

20M train tokens + 1M val tokens, packed into 256-token windows.

Training

  • 2 epochs β‰ˆ 2,440 steps, batch 64Γ—256 = 16,384 tokens/step
  • AdamW (Ξ²=(0.9, 0.95), wd=0.01), grad clip 1.0, fp32
  • Hardware: CPU only (2 vCPU)

Results

MicroSupra-1k (original) MicroSupra-10k (this)
Parameters 1,046 10,000
Train data Fineweb-Edu 300M tok Fineweb-Edu + DCLM-Edu 20M tok
Final train loss 6.046 5.679
Val loss (same held-out val.bin) 6.063 5.697
Perplexity (same val) 429.7 298.0

~10Γ— parameters β†’ βˆ’0.37 nats, ~31% lower perplexity. Still a bacteria model 🦠 β€” it doesn't know facts, but scaling laws are real.

Examples

Prompt: "My name is "
Output: *"My name is .

ASSa, thels

..."*

Prompt: "Question: What is the capital of France?\nAnswer: "
Output: *"Question: What is the capital of France? Answer: .15

S,e.

ec Thet..."*

Usage πŸš€

from huggingface_hub import hf_hub_download
from transformers import LlamaForCausalLM, PreTrainedTokenizerFast
import torch

tokenizer = PreTrainedTokenizerFast(
    tokenizer_file=hf_hub_download("DedeProGames/MicroSupra-10k", "tokenizer.json"),
    bos_token="<s>", eos_token="</s>", pad_token="<pad>", unk_token="<unk>",
)
model = LlamaForCausalLM.from_pretrained("DedeProGames/MicroSupra-10k").eval()

inputs = tokenizer("Question: What is the capital of France?\nAnswer: ", return_tensors="pt")
with torch.no_grad():
    out = model.generate(**inputs, max_new_tokens=100, do_sample=True,
                         temperature=0.35, top_p=0.85, repetition_penalty=1.2,
                         pad_token_id=tokenizer.pad_token_id, eos_token_id=tokenizer.eos_token_id)
print(tokenizer.decode(out[0], skip_special_tokens=True))

Reproducibility

Full pipeline in scripts/:

pip install torch --index-url https://download.pytorch.org/whl/cpu transformers datasets
python scripts/prepare_data.py   # streams both datasets, builds data/{train,val}.bin
python scripts/train.py          # asserts exactly 10,000 params, trains, saves to out/
python scripts/inference.py      # demo generations

Final thoughts

Trained where the GPU can't reach: pure CPU, a 41KB safetensors file, and a model that still can't talk β€” but now with 10Γ— more neurons to not talk with πŸ€–πŸ«Ά

Downloads last month
234
Safetensors
Model size
10k params
Tensor type
F32
Β·
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Datasets used to train DedeProGames/MicroSupra-10k