GoLLeM-v5 128M (final)

Final 128M model of the GoLLeM-v5 series; the best eff of the series.

Author: Arkadiusz Słota / Fabryka AI. Part of the GoLLeM-v5 series of small English language models trained from scratch. All checkpoints of the series, including intermediate ones, are archived in SlayerLab/gollem-v5-ckpts.

Model

  • 122.8M parameters; 16 layers, d_model 768, 12 heads (head dim 64); context 1024 tokens.
  • Architecture: Qwen3-style decoder: RoPE (theta 100,000), SwiGLU (FFN multiplier 2.667), RMSNorm, QK-norm, value residual; tied input/output embeddings.
  • Tokenizer: BPE, 12,288 tokens (tokenizer.json, sha256 3733307577230bb4802d2c774d8e2323f7f64e4139c13712736a57daf91bdda1), 3.8605 bytes per token on WikiText-2 test.
  • Weights: model.safetensors (its sha256 is shown by the Hub for the LFS file), converted tensor by tensor from final_128m_16x768/ckpt_760k.pt (sha256 95d8b43fcd9c5fa06566da13dd03856b00300bad6f37fb4dc9442c18be9ac372) without changing any weight. Only weights are published here: no pickled checkpoint, optimizer state or training logs.
  • Training finished 2026-09-28 13:35 UTC.

Training

  • Recipe: Muon (hidden 2-D weights, lr 0.02) + AdamW (peak lr 6e-4), 2,000 warmup steps, cosine decay to 6e-5, batch 32 x 1024 tokens, 760,000 steps = 24.9B tokens (about 2.65 passes over the pool; each pass uses a new permutation of training windows), seed 1337. Recipe and data were fixed before the run; the data choice was made by a pre-registered comparison against FineWeb-Edu at 32M (a tie, so ARC-MIX was kept).
  • Data: ARC-MIX, 9.39B BPE tokens after a second scan against WikiText-2 and ARC test/validation (normalized 13-gram and short-question matching; matching documents removed) and removal of documents carrying the marker 'CC BY-NC-SA'. Base mixture: SlayerLab/minimal-en-corpus-5b (FineWeb-Edu, DCLM, StackExchange, open-web-math, FineMath, scientific papers, books/Gutenberg, code, CC-News) plus FineWeb-Edu and OpenStax textbooks, with ARC-relevant content upweighted.
  • Result checkpoint: The result is the final checkpoint (step 760,000). A rule written before the end of training would have used the average of the last three checkpoints only if it gained more than 0.3 eff on our selection sets; it gained +0.255, so the last checkpoint is the result.
  • Training run (loss curves, metrics): https://track.fabryka.ai/run/d96091ad-922b-404c-8e8d-6ddaa5fa090a

Evaluation (our measurement, not official)

Measured by us on the result checkpoint with the Glint Tiny-ML board protocol (BLiMP: 67 configs, 67,000 pairs, sentences clipped to 256 tokens, raw log-prob preference; ARC-Easy: test split, zero-shot, LL(question + choice) - LL(question); WikiText-2: byte-normalized perplexity). These are not official scores.

Benchmark Score
ARC-Easy (acc, %) 53.24
BLiMP (acc, %) 79.09
WikiText-2 byte perplexity (lower is better) 2.2538
eff (board formula, includes a size multiplier) 76.94
MultiBLiMP-pl (acc, %, 3,272 pairs) 63.57
  • eff = mean of BLiMP, ARC-Easy and normalized WikiText-2, times a size multiplier that is larger for smaller models; compare eff only between models under the same board formula.
  • This is an English model. Its Polish MultiBLiMP score is close to the length baseline (always choosing the shorter sentence gives 59.67 %); it is reported for completeness, not as Polish ability.
  • Board: https://track.fabryka.ai/models (row "GoLLeM-v5-128M-final-arcmix-16x768").
  • Source of the numbers: liczby_v5.json (sha256 35c9022ee9dbd96d588093c6fc69604d0eaabd18843458671205db41e234dde8), built from the anonymous board read-out and our PL measurement files.

Benchmark overlap disclosure

An earlier scan of ARC-MIX found that 7 of 62 WikiText-2 test articles appeared almost completely in it (as web copies) and 30 of 2,376 ARC-Easy test questions had a match (mostly short factual sentences). For this model the corpus was scanned again before training and all matching documents were removed. BLiMP was covered only by the build-time 13-gram filter.

How to use

The weights are a plain PyTorch state dict saved as safetensors; they are not a transformers AutoModel. The model class is GPT in train_gpt_ref.py in SlayerLab/gollem-v5-ckpts; build it from config.json and load the state dict, then trim logits to 12,288.

Limitations

  • Small base model trained from scratch for research: not instruction-tuned, limited factual knowledge and coherence, not for production use.
  • English only.
  • Single seed; scores come from one checkpoint and our own implementation of the board protocol.

Training data

Base mixture: SlayerLab/minimal-en-corpus-5b (research-mix-5b, 5.0B tokens of its own tokenizer). It is an aggregate corpus: no unified licence is applied and each source keeps its upstream terms. Sources as named in its manifests/mixture.json:

source tokens (M) upstream licence
fineweb-edu 1,101 ODC-By 1.0
dclm-baseline 801 CC BY 4.0
starcoderdata 469 The Stack terms (per-file licences, opt-out)
stackexchange 447 CC BY-SA
loc-pd-books 402 CC0 1.0
wikipedia 362 CC BY-SA 3.0 / GFDL
open-web-math 228 ODC-By
finemath 221 ODC-By
project-gutenberg 210 public domain in the US
scientific-papers 206 unknown
ultrachat 200 MIT; text generated by OpenAI models
wildchat 151 ODC-By; text from conversations with OpenAI models
cc-news 150 unknown (news articles under copyright)
tiny-textbooks 51 Apache 2.0; model-generated text
open-subtitles 1 film subtitles, under copyright
  • Upstream repositories are not recorded in the corpus manifest; licences are those of the Hub datasets matching these source names (not confirmed against the build).
  • Added on top of the base mixture for this series: FineWeb-Edu (ODC-By 1.0) and OpenStax textbooks (CC BY 4.0).
  • Aggregate corpus; each source keeps its upstream terms. Commercial use: review upstream terms, in particular CC-News, scientific papers, StarCoderData and model-generated chat data.

Licence and attribution

  • Weights: no single licence is claimed for the weights yet; they were trained on data under the licences listed in the Training data section above, and any use must respect those terms (license id other, name gollem-v5-see-training-data).
  • OpenStax textbooks: CC BY 4.0, © Rice University, https://openstax.org; titles and editions in OPENSTAX_ATTRIBUTION.md in this repository.
Downloads last month
54
Safetensors
Model size
0.1B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Dataset used to train Maggio33/GoLLeM-v5-128M