Slayer139 1.01b

A small decoder-only English language model (139.3M parameters) trained from scratch by Fabryka AI. It is the alternative version of Slayer139 1.01, trained at the same time with the same recipe and data order on a smaller machine (batch 32 instead of 256). Unlike 1.01, this model is not fine-tuned on ARC.

Author: Arkadiusz Słota (Fabryka AI / SlayerLab). Training, evaluation and safety gates were run by the Fabryka AI agent team (Kolektyw) under his direction.

Result

lm-evaluation-harness 0.4.13, board protocol (C2), our measurement on one AMD Radeon GPU. For comparison, the pretrained base of Slayer139 1.01 (8× RTX 5090, batch 256, same recipe, same tokens, before its ARC fine-tune) on the same machine; our two evaluation machines (RTX 5090, AMD Radeon) gave identical per-item results on this base checkpoint.

measure Slayer139 1.01b 95% CI 1.01 pretrained base 95% CI Δ (1.01 base − 1.01b, paired 95% CI)
BLiMP (67 tasks) 79.64 [79.37; 79.92] 79.41 [79.14; 79.69] −0.23
ARC-Easy (test, 2,376) 54.71 [52.74; 56.69] 57.58 [55.60; 59.55] +2.86
WikiText-2 score 99.49 99.86 +0.37
Overall 77.95 [77.28; 78.62] 78.95 [78.28; 79.62] +1.00 [+0.40; +1.62]

Overall is the mean of the three measures. Confidence intervals: item-level bootstrap (ARC-Easy and each of the 67 BLiMP tasks resampled separately, 10,000 draws); the Δ column uses the same resampled items for both models (paired). They cover test-set sampling only, not seed-to-seed variation; each run is a single seed.

At equal tokens the 8-GPU run (1.01 base) was also lower on held-out DCLM bits-per-byte at the end of training: 0.8651 vs 0.8901 for 1.01b. A rule written before these results ("same quality": ΔDCLM ≤ +0.006 and ΔOverall ≥ −0.3 for the 8-GPU run) is satisfied on both parts.

Provenance (sha256 prefixes of the lm-eval results.json): Slayer139 1.01b a96c8832, 1.01 base 719bd11c; per-item samples d7937049 / 00db3761.

Numbers are our measurements with the board's protocol, not official scores.

Model

  • Architecture: 17 layers × 768, 12 heads (12 KV heads), SwiGLU, decoder-only; vocabulary 24,576 (BPE), context 1,024.
  • Parameters: 139,279,760 unique parameters (input and output embeddings tied: one shared tensor, counted once; counting the shared tensor twice gives 158,154,128).
  • Checkpoint: the final step 722,370 of the run (sha256 418cb898…), chosen as the end of the run by a rule written before any test number.

Training

  • Data: ~23.67B tokens, no repeated epoch, the same data and order as Slayer139 1.01 (sources and licences below). First the ARC-MIX pool, then a top-up from the same source families in the base corpus's proportions, mixed at document level. Both stages were scanned against WikiText-2 (test, validation) and ARC-Easy/Challenge (validation, test) for shared normalized 13-grams, with short ARC questions matched exactly, and against BLiMP sentences of ≥ 6 words; matching documents were removed. ARC train was not filtered.
  • Optimiser and schedule: Muon (hidden matrices, LR 0.02) + AdamW (other parameters, LR 6e-4); WSD schedule: constant to step 577,896, then 1−√ decay to 10% of peak at step 722,370; batch 32 × 1,024 tokens; seed 1337.
  • Hardware: one NVIDIA RTX 5090 up to step 229,845, then 4× RTX 5090 from that checkpoint (data parallel, bf16 all-reduce), ~191k tokens/s on 4 GPUs. The run was resumed from checkpoints several times (including the move to 4 GPUs and two interruptions on 2026-10-08), each time continuing the identical data order. Re-running a stretch of training on 4 GPUs from the same checkpoint changed the training loss by as much as the 1→4 GPU switch did (mean difference ~1e-5), i.e. the switch was indistinguishable from the nondeterminism of repeating the same steps.
  • Training curves: https://track.fabryka.ai/run/ca4d64c7-9485-4c06-a05b-fef81043e1b3

Training data

Aggregate corpus; each source keeps its upstream terms.

  • ARC-MIX pool (SlayerLab/gollem-v5-arcmix-9b, 9.39B tokens before retokenization): the 15-source base SlayerLab/minimal-en-corpus-5b (FineWeb-Edu, DCLM-baseline, open-web-math, FineMath, StarCoderData, StackExchange, LoC public-domain books, Wikipedia, Project Gutenberg, scientific papers, UltraChat, WildChat, CC-News, tiny-textbooks, OpenSubtitles) + a FineWeb-Edu expansion + OpenStax textbooks (×4) + extra copies of FineWeb-Edu documents scored as related to ARC science topics (+2 copies for the highest score, +1 for the next). Removed before training: 4,362 documents by the WikiText-2 / ARC validation+test scan and 1,175 documents marked CC BY-NC-SA.
  • Top-up (11.4M documents, 57.45 GB of text, build manifest b73f45f5…): FineWeb-Edu expansion (37.1%, ODC-By 1.0, plus the same dose of ARC-related copies as the pool), FineWeb-Edu (14.9%, ODC-By 1.0), DCLM-baseline (10.9%, CC BY 4.0), StackExchange via RedPajama (6.1%, CC BY-SA content), LoC public-domain books (5.4%, CC0), StarCoderData (5.1%, The Stack terms), Wikipedia 20231101.en (4.6%, CC BY-SA 3.0 / GFDL), Project Gutenberg (2.9%, US public domain), PMC Open Access commercial-use subset (2.8%, CC0 / CC BY / CC BY-SA per article; ND and NC excluded), open-web-math (2.7%, ODC-By), FineMath (2.6%, ODC-By), WildChat-1M (2.0%, ODC-By; English, non-toxic), CC-News (1.4%, licence unknown), UltraChat 200k (0.7%, MIT), tiny-textbooks (0.7%, Apache 2.0). The top-up scan removed 25,808 documents (0.23%), each together with its ARC-related copies.
  • Licences upstream include ODC-By, CC BY 4.0, CC BY-SA (Wikipedia, StackExchange, some PMC articles), CC0 / public domain, and sources with unknown or restrictive terms. Commercial use: review upstream terms, in particular CC-News, scientific papers, StarCoderData and model-generated chat data (UltraChat, WildChat, tiny-textbooks).
  • OpenStax textbooks (55 titles) are used under CC BY 4.0; attribution: OpenStax, Rice University (see OpenStax attribution below).

Limitations

  • English only; small model; one seed.
  • The pretraining mix deliberately upweights web documents related to ARC science topics (decontaminated against the ARC test set); ARC-Easy is therefore an in-domain benchmark for this model.
  • Not instruction-tuned and not fine-tuned on any benchmark.

OpenStax attribution

The training data of this model includes text extracted from the following OpenStax textbooks, each licensed under the Creative Commons Attribution 4.0 International License (CC BY 4.0), © Rice University. Download for free at https://openstax.org. The texts were obtained from the Hugging Face dataset crumb/openstax-text (revision 8f502ca45f9f05cb5673eae445b7a97a4e8c4349). Modified: text extracted, chunked and filtered (chunks matching benchmark test/validation sets were removed).

Title (from source file name) © year
APBiology 2018
APCollege Physics 2017
APMacroeconomics 2e 2017
APMicroeconomics 2e 2017
Algebra and Trigonometry 2e 2021
American Government 3e 2021
Anatomy and Physiology 2e 2022
Anatomyand Physiology 2017
Astronomy 2e 2022
Astronomy 2018
Biology 2e 2020
Business Ethics 2018
Chemistry 2e 2019
Chemistry Atoms First 2e 2019
College Algebra 2e 2021
College Algebra Corequisite Support 2e 2021
College Physics 2020
College Physics 2e 2022
College Physics for AP Courses 2e 2022
College Success 2020
College Success 2023
Concepts Biology 2017
Contemporary Mathematics 2023
Economics 2e 2018
Economics 3e 2022
Elementary Algebra 2e 2020
Entrepreneurship 2020
Intermediate Algebra 2e 2020
Introduction to Intellectual Property n/a
Introduction to Philosophy 2022
Introduction to Political Science 2022
Introductionto Anthropology 2022
Introductionto Sociology 3e 2021
Introductory Business Statistics 2018
Introductory Statistics 2018
Macroeconomics 2e 2018
Macroeconomics 3e 2022
Microbiology 2021
Microeconomics 2e 2018
Microeconomics 3e 2022
Physics n/a
Prealgebra 2e 2020
Precalculus 2e 2021
Preparing for College Success 2023
Principles Marketing 2023
Principlesof Finance 2022
Psychology 2e 2020
Statistics n/a
USHistory 2021
University Physics Vol 1 2021
University Physics Volume 2 2021
University Physics Volume 3 2021
World History Volume 1 2023
World History Volume 2 2022
Writing Guide 2021

Titles licensed CC BY-NC-SA 4.0, non-English titles, and three CC BY 4.0 titles whose text contains elements marked ‘CC BY-NC-SA’ (Introduction to Business, Organizational Behavior, Principles of Management) were not used.

Reproduce

Evaluation: lm-evaluation-harness 0.4.13, tasks blimp, arc_easy, wikitext (board protocol C2); per-item samples are kept for paired bootstrap. Method and full run history: paper „Same recipe, 4.4× faster on 8× RTX 5090” (Fabryka AI, in preparation).

Downloads last month
19
Safetensors
Model size
0.1B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Datasets used to train SlayerLab/Slayer139-1.01b