GoLLeM v6 — 250M (Polish–English research base model)
STATUS: TRAINING COMPLETE — RESEARCH BASE MODEL. Final checkpoint of a single 760,000-step run. This is a base (pretrained) model, not instruction-tuned: it continues text and does not follow commands. An instruction-tuned version is a separate model with its own card. Results below use our protocol, not official benchmark submissions, and are compared only with our own earlier models (GoLLeM v4 and v5) measured the same way.
GoLLeM v6 is a 252.8M-parameter decoder-only language model trained from scratch on a Polish–English mix (about two thirds Polish, one third English). It is a Fabryka AI project (formerly SlayerLab).
Author: Arkadiusz Słota (Fabryka AI).
What it is for (and what it is not)
Intended use: research on small bilingual language models, comparison with GoLLeM v4 (Polish only) and v5 (English only), and text continuation in Polish and English.
Not intended for: production use, question answering without context, or following instructions. Like any base model of this size it has limited factual knowledge and may produce fluent but wrong text.
Results (final checkpoint, step 760,000)
All numbers: final checkpoint 221b1286…, single run, seed 1337. Our protocol, not official.
English — our implementation of the Glint leaderboard metrics (the same harness we used for GoLLeM v5):
| GoLLeM v6 250M | GoLLeM v5 128M (reference) | |
|---|---|---|
| ARC-Easy (test, 2,376 questions) | 50.21 | 53.24 |
| BLiMP (67,000 pairs) | 77.61 | 79.09 |
| WikiText-2 test, byte perplexity | 2.2884 | 2.2538 |
| eff (with size multiplier) | 74.71 (× 1.0) | 76.94 (76.30 × 1.0084) |
eff = mean(ARC-Easy, BLiMP, WikiScore) × size multiplier, WikiScore = 100·(ln 500 − ln byte_ppl)/(ln 500 − ln 1.86), clipped to 0–100; the multiplier follows the leaderboard formula (1.0 at 150M parameters and above, 1.0084 for v5's declared 122.8M). Without the multiplier the English gap between v6 and v5 128M is −1.59 points, with it −2.23. v6 is below v5 128M in English: v6 saw about 8.9B English tokens (35.8 % of 24.9B), v5 128M was trained on English only.
Polish — our eff-PL protocol (definition fixed on 2026-09-29, before training started):
| GoLLeM v6 250M | GoLLeM v4 250M (reference) | |
|---|---|---|
| G — MultiBLiMP, Polish (3,272 pairs) | 98.14 | 98.47 |
| K — ARC-Easy-PL (EU20 translation, test, 2,376 questions) | 42.09 | 41.96 |
| W — byte perplexity on 463 Polish Wikipedia articles | 1.9578 | 1.9507 |
| eff-PL = mean(G, K, WikiScore on W) | 79.77 | 79.86 |
| PL clean = mean(G on 944 pairs, K on 2,362 questions) | 69.42 ± 0.58 | 69.62 |
- PL clean removes MultiBLiMP pairs and ARC-Easy-PL questions found in GoLLeM v4's training corpus, so that v4 and v6 are compared on the same items: 2,207 of the 3,272 MultiBLiMP pairs are flagged by our scan of the v4 corpus, and pairs with a sentence shorter than 20 characters are also excluded (944 pairs remain); 14 ARC-Easy-PL test questions are in the v4 corpus (2,362 remain). "Clean" is therefore defined relative to the v4 corpus, not the v6 pool. The v6 pool was filtered before training against a protected list that includes these MultiBLiMP sentences (738 documents removed); a separate scan of the finished v6 pool against all 3,272 pairs has not been run. ± is one standard error (binomial, per axis, combined).
- Condition fixed in advance: PL clean ≥ v4 − 1 point. The condition and the two subsets were fixed on 2026-09-29, before any v6 checkpoint was evaluated in Polish (the subsets about one hour after training started); the threshold value, 68.62, depends only on the v4 measurement. The final checkpoint reaches 69.42, so the condition is met.
- Reading: in Polish, v6 is within one standard error of v4 (eff-PL −0.09, PL clean −0.20), while also being trained on about one third English. We do not claim it is better than v4 in Polish.
- W articles were selected by a fixed hash rule and removed if more than 5 % of their 13-grams appear in the v6 training pool or if they were on our privacy drop list (589 of 1,052 removed; 463 kept).
Training progress (evaluated every 6 data slots during training; intermediate checkpoints, not results):
| step | ARC-Easy | BLiMP | WikiText-2 byte_ppl | G | K | W byte_ppl |
|---|---|---|---|---|---|---|
| 60,000 | 43.27 | 73.80 | 2.5193 | 96.49 | 35.82 | 2.1560 |
| 180,000 | 45.29 | 75.54 | 2.4281 | 96.61 | 39.98 | 2.0787 |
| 300,000 | 44.91 | 76.60 | 2.3996 | 97.31 | 39.44 | 2.0514 |
| 420,000 | 47.14 | 76.55 | 2.3922 | 97.04 | 39.23 | 2.0368 |
| 540,000 | 47.56 | 74.99 | 2.3709 | 97.46 | 40.49 | 2.0273 |
| 600,000 (kept trunk checkpoint) | 48.86 | 76.59 | 2.3610 | 97.19 | 40.91 | 2.0239 |
| 720,000 | 49.20 | 77.54 | 2.3069 | 97.83 | 41.25 | 1.9715 |
| 760,000 (final) | 50.21 | 77.61 | 2.2884 | 98.14 | 42.09 | 1.9578 |
The final checkpoint is the best of all 13 evaluated points on every axis; most of the late gain came during the learning-rate decay (last 20 % of steps). With a single run we cannot separate the effect of the decay from the effect of more training.
Training
| Parameters | 252,760,019 |
| Shape | 20 layers, d_model 960, 15 heads, context 1,024 tokens |
| Block | RMSNorm, RoPE (θ = 100,000), SwiGLU (×2.667), QK-norm, value residuals |
| Tokenizer | 32,768-token BPE, Polish + English |
| Optimizer | Muon for the 221,184,000 parameters in 2-D hidden matrices (attention QKV and output projection, SwiGLU gate/up/down): momentum 0.95 (Nesterov), 5 Newton–Schulz iterations, learning rate 0.02 on the same schedule, no weight decay. AdamW (β 0.9 / 0.95, weight decay 0.1, learning rate 6e-4 → 6e-5) for the other 31,576,019: the token embedding tied with the output head, normalisation and QK-norm gains, attention biases and value-residual gates. 2,000 warmup steps. |
| Schedule | warmup–stable–decay: constant learning rate, decay over the last 20 % of steps (from step 608,000) |
| Batch / steps / tokens | 32 × 1,024 tokens per step × 760,000 steps = 24.90B tokens (about one pass over the pool) |
| Validation loss (per token, our held-out set) | 2.5745 at step 760,000 |
| Precision | bf16 autocast with fp32 weights; torch.compile. Released weights are fp32. |
| Hardware / time | 1 × NVIDIA RTX PRO 6000 Blackwell Server Edition (cloud), single GPU; about 99,000 tokens/s during training. 2026-09-29 18:23 → 2026-10-03 00:19 UTC (about 78 h wall-clock, including a pause after each data slot for the control check: median about 6 min, range 5–35 min, about 8 h in total). |
Data queue with a rule. Training data arrived in 76 slots of 10,000 steps. After each slot the checkpoint was scored on five fixed control sets (English educational, encyclopedic and Q&A text; Polish encyclopedic and general text) on a separate machine. The first 6 slots ran in shadow mode: the rule was not applied, and its noise level was calibrated on them (per-set residual standard deviation σ, at least 0.002), so the rule's thresholds were set during the run, after the shadow slots. From slot 6 on, a slot would be rejected if the loss on any of the three English control sets rose by more than 3σ; the two Polish sets were scored and logged but did not decide. All 76 slots were accepted, none rejected.
Training data
GoLLeM v6 250M was pretrained on a fixed pool of 25,000,424,215 tokens (26,124,390 documents; tokenizer sha256
0640d3bd…); each token was seen at most once (24.90B of the 25.00B tokens were used). The pool is an aggregate of public
sources; each source keeps its upstream terms.
| part | share | tokens | origin | upstream license |
|---|---|---|---|---|
| English mix (ARC-MIX v5) | 35.79 % | 8.95 B | FineWeb-Edu, DCLM-baseline, FineMath, OpenWebMath, Wikipedia (EN), StackExchange, StarCoderData, LoC PD Books, Project Gutenberg, scientific papers, CC-News, UltraChat, WildChat, tiny-textbooks, OpenSubtitles (small), OpenStax textbooks | mixed: ODC-By, CC BY 4.0, CC BY-SA, CC0/PD, The Stack terms, unknown (CC-News, scientific papers), model-generated (UltraChat, WildChat, tiny-textbooks) |
| Polish web (HPLT 3.0, filtered) | 60.20 % | 15.05 B | HPLT/HPLT3.0, Polish, our quality filtering | compilation CC0-1.0; texts collected under TDM exception (EU DSM art. 4) with crawl-level opt-out |
| Polish Wikipedia + Wolne Lektury (SA/FAL) | 1.94 % | 0.49 B | SlayerLab/polish-dynaword | CC BY-SA 3.0, CC BY-SA 3.0/4.0, Free Art License 1.3 |
| Polish open: Europarl v7, Wolne Lektury (PD) | 0.33 % | 0.08 B | statmt.org Europarl v7, polish-dynaword | Europarl: no known copyright restrictions; public domain |
| Biblioteka Nauki (CC BY-SA), e-mails masked | 0.60 % | 0.15 B | polish-dynaword biblioteka_nauki |
CC BY-SA 4.0 |
| Biblioteka Nauki (CC BY / PD), e-mails masked | 1.14 % | 0.28 B | polish-dynaword biblioteka_nauki |
CC BY 4.0, public domain |
- No non-commercial (NC) sources. Documents marked CC BY-NC-SA were removed from the English mix.
- Share-alike sources are about 7 % of tokens. The model weights are released under CC BY-SA 4.0.
- Commercial use: review upstream terms, in particular CC-News, scientific papers, StarCoderData (The Stack terms) and model-generated chat data (UltraChat, WildChat, tiny-textbooks: outputs of OpenAI models).
- Attribution: the English mix includes text from 55 OpenStax textbooks (CC BY 4.0, © Rice University,
https://openstax.org), obtained from
crumb/openstax-text(revision8f502ca4…), modified (extracted, chunked, filtered). The full title list is in the appendix OpenStax attribution below, because the English mix itself is not published. - Personal data: in Biblioteka Nauki, e-mail addresses were masked before training (
[email]). - Benchmark decontamination: documents matching WikiText-2 test/validation or ARC validation/test were removed from the English mix (4,362 documents). In the Polish part, documents containing an exact copy (after normalization) of any of 12,755 protected Polish evaluation sentences of at least 20 characters, including MultiBLiMP-pl, were dropped (738 documents).
Data availability
| part | status |
|---|---|
| upstream sources | public (links above) |
| Polish Wikipedia, Wolne Lektury, Biblioteka Nauki | the exact input files are public in SlayerLab/polish-dynaword (sha256 match); our e-mail masking of Biblioteka Nauki is not published |
| Polish web (HPLT, filtered) | not published in the exact form used (18 files); upstream HPLT 3.0 is public |
| English mix (ARC-MIX v5) | not published as data; recipe described in the SlayerLab/gollem-v5-arcmix-9b card |
| tokenized pool | not published |
Usage
The model is a custom PyTorch architecture, not a transformers class. The repository contains:
| file | content | sha256 |
|---|---|---|
model.safetensors |
fp32 weights (252,760,019 parameters; output head stored as a copy of the tied embedding) | d30b60d5f617a1d118e4ba363996c5e56ea57037b82dbb707382a0be29e7bbd0 |
config.json |
architecture | 61d8594034918d7e49253f3801da79bae9f14f57ba871fccde6d307a01c73a8f |
tokenizer.json |
32,768-token BPE (tokenizers) |
0640d3bd3674d7a6e59c540d945a2ad275a12e1bcaeb6007c37a90751f9815dd |
modeling_gollem_v6.py |
model code (inference) and a small generate helper |
ee2ea8c4e957413f3a7c1e8fe93a0a0b4431a2bcabaa6cc5ceeaea103fcd58ee |
# pip install torch safetensors tokenizers huggingface_hub
from huggingface_hub import snapshot_download
import sys
path = snapshot_download("SlayerLab/GoLLeM-v6-250M") # pin revision="<commit>" for reproducibility
sys.path.insert(0, path)
from modeling_gollem_v6 import load_gollem_v6, generate
model, tok = load_gollem_v6(path) # strict load, eval mode, CPU by default
print(generate(model, tok, "Ala ma kota", max_new_tokens=40, temperature=0.8, top_k=50, seed=1))
print(generate(model, tok, "Photosynthesis is the process by which", max_new_tokens=40))
- This is a base model: it continues text; it does not answer questions or follow instructions.
- Context: 1,024 tokens. RoPE uses the interleaved channel convention; if you port the model to another framework, keep that convention.
- The module definitions in
modeling_gollem_v6.pyare those used in training (initialisation and training loss removed). Loadingmodel.safetensorsinto them reproduces the logits of the training checkpoint exactly (checked withtorch.equal, also on the files downloaded from this repository).
Limitations
- Small base model: not instruction-tuned, limited factual knowledge, may reproduce biases and errors present in web data. Not for production use.
- Single run, single seed. Differences below about one standard error (or below the run-to-run noise we measured on smaller models, about 0.3 eff) should not be read as real.
- English and Polish scores use our protocol; they are not official leaderboard results, and self-reported numbers of other models measured differently are not comparable.
- Self-description. Prompted in chat format (ChatML), the model may describe itself as "an AI language model created by OpenAI". The English pretraining mix contains model-generated chat data (UltraChat, WildChat) with such self-descriptions. GoLLeM is not affiliated with OpenAI; it was built by Fabryka AI (author: Arkadiusz Słota).
License
Weights: CC BY-SA 4.0. Attribution: GoLLeM v6, Arkadiusz Słota / Fabryka AI, link to this repository; derivative weights under the same licence.
Training data keep their upstream licences; see Training data above.
Po polsku (skrót)
GoLLeM v6 250M to bazowy (niedostrojony do poleceń) model językowy polsko-angielski, 252,8 mln parametrów, wytrenowany od zera na 24,9 mld tokenów (ok. 2/3 polski, 1/3 angielski). Wyniki według naszego protokołu, nie oficjalne: eff (angielski) 74,71, eff-PL 79,77, PL czyste 69,42 ± 0,58. Warunek ustalony przed pierwszym pomiarem po polsku (PL czyste ≥ 68,62, czyli v4 − 1 pkt) jest spełniony. W polskim v6 jest w granicach jednego błędu standardowego od GoLLeM v4 (który trenowano tylko po polsku), w angielskim poniżej GoLLeM v5 128M (tylko angielski). Model do badań, nie do zastosowań produkcyjnych. W formacie czatu (ChatML) model może przedstawiać się jako „model językowy stworzony przez OpenAI”, bo angielska część danych zawiera rozmowy wygenerowane modelami OpenAI (UltraChat, WildChat). GoLLeM nie jest powiązany z OpenAI; stworzyła go Fabryka AI (autor: Arkadiusz Słota).
Appendix: OpenStax attribution
The training data of this model includes text extracted from the following OpenStax textbooks, each licensed under the Creative Commons Attribution 4.0 International License (CC BY 4.0), © Rice University. Download for free at https://openstax.org. The texts were obtained from the Hugging Face dataset crumb/openstax-text (revision 8f502ca45f9f05cb5673eae445b7a97a4e8c4349). Modified: text extracted, chunked and filtered (chunks matching benchmark test/validation sets were removed).
| Title (from source file name) | © year |
|---|---|
| APBiology | 2018 |
| APCollege Physics | 2017 |
| APMacroeconomics 2e | 2017 |
| APMicroeconomics 2e | 2017 |
| Algebra and Trigonometry 2e | 2021 |
| American Government 3e | 2021 |
| Anatomy and Physiology 2e | 2022 |
| Anatomyand Physiology | 2017 |
| Astronomy 2e | 2022 |
| Astronomy | 2018 |
| Biology 2e | 2020 |
| Business Ethics | 2018 |
| Chemistry 2e | 2019 |
| Chemistry Atoms First 2e | 2019 |
| College Algebra 2e | 2021 |
| College Algebra Corequisite Support 2e | 2021 |
| College Physics | 2020 |
| College Physics 2e | 2022 |
| College Physics for AP Courses 2e | 2022 |
| College Success | 2020 |
| College Success | 2023 |
| Concepts Biology | 2017 |
| Contemporary Mathematics | 2023 |
| Economics 2e | 2018 |
| Economics 3e | 2022 |
| Elementary Algebra 2e | 2020 |
| Entrepreneurship | 2020 |
| Intermediate Algebra 2e | 2020 |
| Introduction to Intellectual Property | n/a |
| Introduction to Philosophy | 2022 |
| Introduction to Political Science | 2022 |
| Introductionto Anthropology | 2022 |
| Introductionto Sociology 3e | 2021 |
| Introductory Business Statistics | 2018 |
| Introductory Statistics | 2018 |
| Macroeconomics 2e | 2018 |
| Macroeconomics 3e | 2022 |
| Microbiology | 2021 |
| Microeconomics 2e | 2018 |
| Microeconomics 3e | 2022 |
| Physics | n/a |
| Prealgebra 2e | 2020 |
| Precalculus 2e | 2021 |
| Preparing for College Success | 2023 |
| Principles Marketing | 2023 |
| Principlesof Finance | 2022 |
| Psychology 2e | 2020 |
| Statistics | n/a |
| USHistory | 2021 |
| University Physics Vol 1 | 2021 |
| University Physics Volume 2 | 2021 |
| University Physics Volume 3 | 2021 |
| World History Volume 1 | 2023 |
| World History Volume 2 | 2022 |
| Writing Guide | 2021 |
Titles licensed CC BY-NC-SA 4.0, non-English titles, and three CC BY 4.0 titles whose text contains elements marked ‘CC BY-NC-SA’ (Introduction to Business, Organizational Behavior, Principles of Management) were not used.
GoLLeM v6 — Fabryka AI. Author: Arkadiusz Słota. Research base model (final checkpoint).
- Downloads last month
- 354