TinyBrainBot-100M-v3-Base

A 100M-parameter (75.5M non-embedding) from-scratch base language model, trained on 25B tokens. Same size and class as SupraLabs/Supra2-100M, and it beats Supra2-100M-Base on 6/7 benchmarks on the official EleutherAI LM-Eval Harness — at a smaller token budget (25B vs 30B).

  • Architecture: Llama-compatible — hidden 768, 12 layers, 12 heads / 4 KV, FFN 2048, context 1024, vocab 32,000, tied embeddings.
  • Training: 25B-token WSD pretrain (fineweb-edu / dclm / wikipedia / gutenberg / cosmopedia + a quality anneal), then a short knowledge-focused continuation.

Benchmarks (EleutherAI lm-eval, 0-shot, acc_norm; WinoGrande/MMLU = acc)

Benchmark This model Supra2-100M-Base Supra2-100M-Instruct
ARC-Easy 54.7 47.8 44.4
ARC-Challenge 30.1 24.8 24.7
OpenBookQA 34.0 32.0 30.4
WinoGrande 53.0 50.7 50.5
PIQA 66.2 65.5 64.4
MMLU 25.0 23.3 25.8
HellaSwag 32.6 36.0 35.9

6/7 vs Supra2-Base and 5/7 vs Supra2-Instruct. HellaSwag is the one benchmark where Supra leads.

Reproduce these numbers

EleutherAI lm-eval-harness v0.4.12, 0-shot, evaluated on the HF repo (do not eval the GGUF — llama.cpp's --multiple-choice path under-reports these tasks; the .hellaswag path is fine):

lm_eval --model hf \
  --model_args pretrained=nkthebass/tinybrainbot-100m-v3-base,dtype=float32 \
  --tasks hellaswag,arc_easy,arc_challenge,openbookqa,winogrande,piqa,mmlu \
  --num_fewshot 0 --batch_size 32

Metrics: acc_norm for HellaSwag / ARC-Easy / ARC-Challenge / OpenBookQA / PIQA; acc for WinoGrande & MMLU.

Usage

from transformers import AutoModelForCausalLM, AutoTokenizer
tok = AutoTokenizer.from_pretrained("nkthebass/tinybrainbot-100m-v3-base")
model = AutoModelForCausalLM.from_pretrained("nkthebass/tinybrainbot-100m-v3-base")
ids = tok("The capital of France is", return_tensors="pt").input_ids
print(tok.decode(model.generate(ids, max_new_tokens=20)[0]))

GGUF

An F16 GGUF is included (tinybrainbot-100m-v3-base-f16.gguf) for llama.cpp / Ollama / LM Studio, with the correct add_space_prefix=false + leading-space template baked in for faithful tokenization.

Limitations

A 100M base model: strong on multiple-choice reasoning for its size, but open-ended generation is limited and can be factually unreliable. For chat use the instruct variant; for arithmetic use the math variant. Not aligned or safety-tuned.

Companion models: instruct · math.

Downloads last month
921
Safetensors
Model size
0.1B params
Tensor type
F16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support