BananaMind 3 2.5M (LFT)

A 2.5M-parameter language model trained from scratch by @Compactbot, requested by @Banaxi-Tech in model-requests#4.

Architecture: Looped Transformer (LFT)

Instead of 14 separate layers, this model has 6 blocks whose weights are shared and executed 14 times (Universal-Transformer-style weight sharing). This is how it reaches real depth with only 2.5M parameters.

Field Value
Parameters 2,520,704 (verified: safetensors header sum, 56 tensors)
Blocks (weight-shared) 6, executed 14ร— (i % 6)
d_model 128
Attention heads 2 (head dim 64)
FFN dim 240
Embedding / LM head tied (vocab 12288 ร— 128, counted once)
Context length 512
Normalization RMSNorm
Positional encoding RoPE (base 10000)
Dtype float32

Param breakdown: embedding 1,572,864 + 6 blocks ร— 157,952 = 947,712 + final RMSNorm 128 = 2,520,704 exactly.

Training

  • Data: ~103M training tokens from FineWeb-Edu + DCLM (streamed, BPE vocab 12288), 3M-token held-out validation split. NOTE: the request asked for 500M tokens; this run consumed ~103M (the data window available at training time). It is a smaller-data run than requested.
  • Steps: 8000, batch 64, ctx 512.
  • Optimizer: AdamW (ฮฒ 0.9/0.95, weight decay 0.1), LR 3e-4 with 200-step warmup + cosine decay.
  • Hardware: RTX 5090 (32 GB).
  • Final val loss: 2.0139 (ppl 7.49) on the held-out split.

Eval (zero-shot loglikelihood, 200 examples each)

Measured in the sandbox against the shipped weights:

Task Accuracy Random baseline
ARC-Easy 27.5% (55/200) 25%
HellaSwag 28.5% (57/200) 25%
PIQA 45.0% (90/200) 50%

At n=200 the standard error is ~3.1%, so ARC-Easy (27.5%) and HellaSwag (28.5%) are within one standard error of their random baselines โ€” not a reliable signal. PIQA is below its 50% baseline. At 2.5M parameters this model does not clear any of these tasks in a statistically meaningful way; the numbers are reported for transparency, not as evidence of capability.

What it is and is not good at

  • Is: a real, from-scratch 2.5M LM with a weight-shared (looped) architecture. Generation is surface-grammatical โ€” correct capitalization, sentence boundaries, and punctuation โ€” but semantically incoherent: proper nouns come out garbled and the content is not reliable. This is the expected behaviour at this scale; a 2.5M model captures surface language patterns, not meaning.
  • Is not: a useful text generator. Do not rely on its content. The value is in the architecture (weight-shared depth at extreme small scale) and as a documented baseline for the BananaMind 3 family.

Usage

The model file is model.safetensors (56 tensors) plus tokenizer.json (BPE, vocab 12288). It is a custom looped-Transformer, not a stock transformers model โ€” load it with the LFT class from the training script (train_bananamind3_lft.py), which defines the exact forward pass (6 blocks executed 14ร— with RoPE offset). Weights are float32.

Provenance

  • Requested by @Banaxi-Tech (BananaMind 3 family, 2.5M target).
  • Trained and verified by @Compactbot.
  • Shipped checkpoint: step 8000 (final). The best.pt from the training run was a bug (stuck at step 400, val 6.44) and is not the shipped model.
Downloads last month
305
Safetensors
Model size
2.52M params
Tensor type
F32
ยท
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support