GoLLeM v6 — 250M (Polish–English research base model)

STATUS: TRAINING COMPLETE — RESEARCH BASE MODEL. Final checkpoint of a single 760,000-step run. This is a base (pretrained) model, not instruction-tuned: it continues text and does not follow commands. An instruction-tuned version is a separate model with its own card. Results below use our protocol, not official benchmark submissions, and are compared only with our own earlier models (GoLLeM v4 and v5) measured the same way.

GoLLeM v6 is a 252.8M-parameter decoder-only language model trained from scratch on a Polish–English mix (about two thirds Polish, one third English). It is a Fabryka AI project (formerly SlayerLab).

Author: Arkadiusz Słota (Fabryka AI).

What it is for (and what it is not)

Intended use: research on small bilingual language models, comparison with GoLLeM v4 (Polish only) and v5 (English only), and text continuation in Polish and English.

Not intended for: production use, question answering without context, or following instructions. Like any base model of this size it has limited factual knowledge and may produce fluent but wrong text.

Results (final checkpoint, step 760,000)

All numbers: final checkpoint 221b1286…, single run, seed 1337. Our protocol, not official.

English — our implementation of the Glint leaderboard metrics (the same harness we used for GoLLeM v5):

GoLLeM v6 250M GoLLeM v5 128M (reference)
ARC-Easy (test, 2,376 questions) 50.21 53.24
BLiMP (67,000 pairs) 77.61 79.09
WikiText-2 test, byte perplexity 2.2884 2.2538
eff (with size multiplier) 74.71 (× 1.0) 76.94 (76.30 × 1.0084)

eff = mean(ARC-Easy, BLiMP, WikiScore) × size multiplier, WikiScore = 100·(ln 500 − ln byte_ppl)/(ln 500 − ln 1.86), clipped to 0–100; the multiplier follows the leaderboard formula (1.0 at 150M parameters and above, 1.0084 for v5's declared 122.8M). Without the multiplier the English gap between v6 and v5 128M is −1.59 points, with it −2.23. v6 is below v5 128M in English: v6 saw about 8.9B English tokens (35.8 % of 24.9B), v5 128M was trained on English only.

Polish — our eff-PL protocol (definition fixed on 2026-09-29, before training started):

GoLLeM v6 250M GoLLeM v4 250M (reference)
G — MultiBLiMP, Polish (3,272 pairs) 98.14 98.47
K — ARC-Easy-PL (EU20 translation, test, 2,376 questions) 42.09 41.96
W — byte perplexity on 463 Polish Wikipedia articles 1.9578 1.9507
eff-PL = mean(G, K, WikiScore on W) 79.77 79.86
PL clean = mean(G on 944 pairs, K on 2,362 questions) 69.42 ± 0.58 69.62
  • PL clean removes MultiBLiMP pairs and ARC-Easy-PL questions found in GoLLeM v4's training corpus, so that v4 and v6 are compared on the same items: 2,207 of the 3,272 MultiBLiMP pairs are flagged by our scan of the v4 corpus, and pairs with a sentence shorter than 20 characters are also excluded (944 pairs remain); 14 ARC-Easy-PL test questions are in the v4 corpus (2,362 remain). "Clean" is therefore defined relative to the v4 corpus, not the v6 pool. The v6 pool was filtered before training against a protected list that includes these MultiBLiMP sentences (738 documents removed); a separate scan of the finished v6 pool against all 3,272 pairs has not been run. ± is one standard error (binomial, per axis, combined).
  • Condition fixed in advance: PL clean ≥ v4 − 1 point. The condition and the two subsets were fixed on 2026-09-29, before any v6 checkpoint was evaluated in Polish (the subsets about one hour after training started); the threshold value, 68.62, depends only on the v4 measurement. The final checkpoint reaches 69.42, so the condition is met.
  • Reading: in Polish, v6 is within one standard error of v4 (eff-PL −0.09, PL clean −0.20), while also being trained on about one third English. We do not claim it is better than v4 in Polish.
  • W articles were selected by a fixed hash rule and removed if more than 5 % of their 13-grams appear in the v6 training pool or if they were on our privacy drop list (589 of 1,052 removed; 463 kept).

Training progress (evaluated every 6 data slots during training; intermediate checkpoints, not results):

step ARC-Easy BLiMP WikiText-2 byte_ppl G K W byte_ppl
60,000 43.27 73.80 2.5193 96.49 35.82 2.1560
180,000 45.29 75.54 2.4281 96.61 39.98 2.0787
300,000 44.91 76.60 2.3996 97.31 39.44 2.0514
420,000 47.14 76.55 2.3922 97.04 39.23 2.0368
540,000 47.56 74.99 2.3709 97.46 40.49 2.0273
600,000 (kept trunk checkpoint) 48.86 76.59 2.3610 97.19 40.91 2.0239
720,000 49.20 77.54 2.3069 97.83 41.25 1.9715
760,000 (final) 50.21 77.61 2.2884 98.14 42.09 1.9578

The final checkpoint is the best of all 13 evaluated points on every axis; most of the late gain came during the learning-rate decay (last 20 % of steps). With a single run we cannot separate the effect of the decay from the effect of more training.

Training

Parameters 252,760,019
Shape 20 layers, d_model 960, 15 heads, context 1,024 tokens
Block RMSNorm, RoPE (θ = 100,000), SwiGLU (×2.667), QK-norm, value residuals
Tokenizer 32,768-token BPE, Polish + English
Optimizer Muon for the 221,184,000 parameters in 2-D hidden matrices (attention QKV and output projection, SwiGLU gate/up/down): momentum 0.95 (Nesterov), 5 Newton–Schulz iterations, learning rate 0.02 on the same schedule, no weight decay. AdamW (β 0.9 / 0.95, weight decay 0.1, learning rate 6e-4 → 6e-5) for the other 31,576,019: the token embedding tied with the output head, normalisation and QK-norm gains, attention biases and value-residual gates. 2,000 warmup steps.
Schedule warmup–stable–decay: constant learning rate, decay over the last 20 % of steps (from step 608,000)
Batch / steps / tokens 32 × 1,024 tokens per step × 760,000 steps = 24.90B tokens (about one pass over the pool)
Validation loss (per token, our held-out set) 2.5745 at step 760,000
Precision bf16 autocast with fp32 weights; torch.compile. Released weights are fp32.
Hardware / time 1 × NVIDIA RTX PRO 6000 Blackwell Server Edition (cloud), single GPU; about 99,000 tokens/s during training. 2026-09-29 18:23 → 2026-10-03 00:19 UTC (about 78 h wall-clock, including a pause after each data slot for the control check: median about 6 min, range 5–35 min, about 8 h in total).

Data queue with a rule. Training data arrived in 76 slots of 10,000 steps. After each slot the checkpoint was scored on five fixed control sets (English educational, encyclopedic and Q&A text; Polish encyclopedic and general text) on a separate machine. The first 6 slots ran in shadow mode: the rule was not applied, and its noise level was calibrated on them (per-set residual standard deviation σ, at least 0.002), so the rule's thresholds were set during the run, after the shadow slots. From slot 6 on, a slot would be rejected if the loss on any of the three English control sets rose by more than 3σ; the two Polish sets were scored and logged but did not decide. All 76 slots were accepted, none rejected.

Training data

GoLLeM v6 250M was pretrained on a fixed pool of 25,000,424,215 tokens (26,124,390 documents; tokenizer sha256 0640d3bd…); each token was seen at most once (24.90B of the 25.00B tokens were used). The pool is an aggregate of public sources; each source keeps its upstream terms.

part share tokens origin upstream license
English mix (ARC-MIX v5) 35.79 % 8.95 B FineWeb-Edu, DCLM-baseline, FineMath, OpenWebMath, Wikipedia (EN), StackExchange, StarCoderData, LoC PD Books, Project Gutenberg, scientific papers, CC-News, UltraChat, WildChat, tiny-textbooks, OpenSubtitles (small), OpenStax textbooks mixed: ODC-By, CC BY 4.0, CC BY-SA, CC0/PD, The Stack terms, unknown (CC-News, scientific papers), model-generated (UltraChat, WildChat, tiny-textbooks)
Polish web (HPLT 3.0, filtered) 60.20 % 15.05 B HPLT/HPLT3.0, Polish, our quality filtering compilation CC0-1.0; texts collected under TDM exception (EU DSM art. 4) with crawl-level opt-out
Polish Wikipedia + Wolne Lektury (SA/FAL) 1.94 % 0.49 B SlayerLab/polish-dynaword CC BY-SA 3.0, CC BY-SA 3.0/4.0, Free Art License 1.3
Polish open: Europarl v7, Wolne Lektury (PD) 0.33 % 0.08 B statmt.org Europarl v7, polish-dynaword Europarl: no known copyright restrictions; public domain
Biblioteka Nauki (CC BY-SA), e-mails masked 0.60 % 0.15 B polish-dynaword biblioteka_nauki CC BY-SA 4.0
Biblioteka Nauki (CC BY / PD), e-mails masked 1.14 % 0.28 B polish-dynaword biblioteka_nauki CC BY 4.0, public domain
  • No non-commercial (NC) sources. Documents marked CC BY-NC-SA were removed from the English mix.
  • Share-alike sources are about 7 % of tokens. The model weights are released under CC BY-SA 4.0.
  • Commercial use: review upstream terms, in particular CC-News, scientific papers, StarCoderData (The Stack terms) and model-generated chat data (UltraChat, WildChat, tiny-textbooks: outputs of OpenAI models).
  • Attribution: the English mix includes text from 55 OpenStax textbooks (CC BY 4.0, © Rice University, https://openstax.org), obtained from crumb/openstax-text (revision 8f502ca4…), modified (extracted, chunked, filtered). The full title list is in the appendix OpenStax attribution below, because the English mix itself is not published.
  • Personal data: in Biblioteka Nauki, e-mail addresses were masked before training ([email]).
  • Benchmark decontamination: documents matching WikiText-2 test/validation or ARC validation/test were removed from the English mix (4,362 documents). In the Polish part, documents containing an exact copy (after normalization) of any of 12,755 protected Polish evaluation sentences of at least 20 characters, including MultiBLiMP-pl, were dropped (738 documents).

Data availability

part status
upstream sources public (links above)
Polish Wikipedia, Wolne Lektury, Biblioteka Nauki the exact input files are public in SlayerLab/polish-dynaword (sha256 match); our e-mail masking of Biblioteka Nauki is not published
Polish web (HPLT, filtered) not published in the exact form used (18 files); upstream HPLT 3.0 is public
English mix (ARC-MIX v5) not published as data; recipe described in the SlayerLab/gollem-v5-arcmix-9b card
tokenized pool not published

Usage

The model is a custom PyTorch architecture, not a transformers class. The repository contains:

file content sha256
model.safetensors fp32 weights (252,760,019 parameters; output head stored as a copy of the tied embedding) d30b60d5f617a1d118e4ba363996c5e56ea57037b82dbb707382a0be29e7bbd0
config.json architecture 61d8594034918d7e49253f3801da79bae9f14f57ba871fccde6d307a01c73a8f
tokenizer.json 32,768-token BPE (tokenizers) 0640d3bd3674d7a6e59c540d945a2ad275a12e1bcaeb6007c37a90751f9815dd
modeling_gollem_v6.py model code (inference) and a small generate helper ee2ea8c4e957413f3a7c1e8fe93a0a0b4431a2bcabaa6cc5ceeaea103fcd58ee
# pip install torch safetensors tokenizers huggingface_hub
from huggingface_hub import snapshot_download
import sys

path = snapshot_download("SlayerLab/GoLLeM-v6-250M")   # pin revision="<commit>" for reproducibility
sys.path.insert(0, path)
from modeling_gollem_v6 import load_gollem_v6, generate

model, tok = load_gollem_v6(path)                 # strict load, eval mode, CPU by default
print(generate(model, tok, "Ala ma kota", max_new_tokens=40, temperature=0.8, top_k=50, seed=1))
print(generate(model, tok, "Photosynthesis is the process by which", max_new_tokens=40))
  • This is a base model: it continues text; it does not answer questions or follow instructions.
  • Context: 1,024 tokens. RoPE uses the interleaved channel convention; if you port the model to another framework, keep that convention.
  • The module definitions in modeling_gollem_v6.py are those used in training (initialisation and training loss removed). Loading model.safetensors into them reproduces the logits of the training checkpoint exactly (checked with torch.equal, also on the files downloaded from this repository).

Limitations

  • Small base model: not instruction-tuned, limited factual knowledge, may reproduce biases and errors present in web data. Not for production use.
  • Single run, single seed. Differences below about one standard error (or below the run-to-run noise we measured on smaller models, about 0.3 eff) should not be read as real.
  • English and Polish scores use our protocol; they are not official leaderboard results, and self-reported numbers of other models measured differently are not comparable.
  • Self-description. Prompted in chat format (ChatML), the model may describe itself as "an AI language model created by OpenAI". The English pretraining mix contains model-generated chat data (UltraChat, WildChat) with such self-descriptions. GoLLeM is not affiliated with OpenAI; it was built by Fabryka AI (author: Arkadiusz Słota).

License

Weights: CC BY-SA 4.0. Attribution: GoLLeM v6, Arkadiusz Słota / Fabryka AI, link to this repository; derivative weights under the same licence.

Training data keep their upstream licences; see Training data above.

Po polsku (skrót)

GoLLeM v6 250M to bazowy (niedostrojony do poleceń) model językowy polsko-angielski, 252,8 mln parametrów, wytrenowany od zera na 24,9 mld tokenów (ok. 2/3 polski, 1/3 angielski). Wyniki według naszego protokołu, nie oficjalne: eff (angielski) 74,71, eff-PL 79,77, PL czyste 69,42 ± 0,58. Warunek ustalony przed pierwszym pomiarem po polsku (PL czyste ≥ 68,62, czyli v4 − 1 pkt) jest spełniony. W polskim v6 jest w granicach jednego błędu standardowego od GoLLeM v4 (który trenowano tylko po polsku), w angielskim poniżej GoLLeM v5 128M (tylko angielski). Model do badań, nie do zastosowań produkcyjnych. W formacie czatu (ChatML) model może przedstawiać się jako „model językowy stworzony przez OpenAI”, bo angielska część danych zawiera rozmowy wygenerowane modelami OpenAI (UltraChat, WildChat). GoLLeM nie jest powiązany z OpenAI; stworzyła go Fabryka AI (autor: Arkadiusz Słota).

Appendix: OpenStax attribution

The training data of this model includes text extracted from the following OpenStax textbooks, each licensed under the Creative Commons Attribution 4.0 International License (CC BY 4.0), © Rice University. Download for free at https://openstax.org. The texts were obtained from the Hugging Face dataset crumb/openstax-text (revision 8f502ca45f9f05cb5673eae445b7a97a4e8c4349). Modified: text extracted, chunked and filtered (chunks matching benchmark test/validation sets were removed).

Title (from source file name) © year
APBiology 2018
APCollege Physics 2017
APMacroeconomics 2e 2017
APMicroeconomics 2e 2017
Algebra and Trigonometry 2e 2021
American Government 3e 2021
Anatomy and Physiology 2e 2022
Anatomyand Physiology 2017
Astronomy 2e 2022
Astronomy 2018
Biology 2e 2020
Business Ethics 2018
Chemistry 2e 2019
Chemistry Atoms First 2e 2019
College Algebra 2e 2021
College Algebra Corequisite Support 2e 2021
College Physics 2020
College Physics 2e 2022
College Physics for AP Courses 2e 2022
College Success 2020
College Success 2023
Concepts Biology 2017
Contemporary Mathematics 2023
Economics 2e 2018
Economics 3e 2022
Elementary Algebra 2e 2020
Entrepreneurship 2020
Intermediate Algebra 2e 2020
Introduction to Intellectual Property n/a
Introduction to Philosophy 2022
Introduction to Political Science 2022
Introductionto Anthropology 2022
Introductionto Sociology 3e 2021
Introductory Business Statistics 2018
Introductory Statistics 2018
Macroeconomics 2e 2018
Macroeconomics 3e 2022
Microbiology 2021
Microeconomics 2e 2018
Microeconomics 3e 2022
Physics n/a
Prealgebra 2e 2020
Precalculus 2e 2021
Preparing for College Success 2023
Principles Marketing 2023
Principlesof Finance 2022
Psychology 2e 2020
Statistics n/a
USHistory 2021
University Physics Vol 1 2021
University Physics Volume 2 2021
University Physics Volume 3 2021
World History Volume 1 2023
World History Volume 2 2022
Writing Guide 2021

Titles licensed CC BY-NC-SA 4.0, non-English titles, and three CC BY 4.0 titles whose text contains elements marked ‘CC BY-NC-SA’ (Introduction to Business, Organizational Behavior, Principles of Management) were not used.


GoLLeM v6 — Fabryka AI. Author: Arkadiusz Słota. Research base model (final checkpoint).

Downloads last month
354
Safetensors
Model size
0.3B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support