TokenizerTest-32m-B (Polish research base model, tokenizer comparison — arm B)
A small Polish base language model trained as arm B of a controlled tokenizer experiment. It is a research checkpoint for comparing tokenizers, not a model for use. It is a Fabryka AI project (formerly SlayerLab).
Author: Arkadiusz Słota (Fabryka AI), who designed and ran the experiment. Tokenizer of this arm: stubbornGuy.
Companion model (tokenizer A/B experiment)
This model is one half of a controlled tokenizer comparison. Two models were trained with the same architecture, training recipe, seed and the same Polish text. The only difference is the tokenizer:
| arm | tokenizer | repo | final step |
|---|---|---|---|
| A | Fabryka AI BPE, vocab 32 000 (60e23148) |
SlayerLab/TokenizerTest-32m-A | 89 000 |
| B | tokenizer by stubbornGuy, vocab 32 000 (77eeebed) |
SlayerLab/TokenizerTest-32m-B | 95 000 |
This model: B.
Shared recipe: ~38.9M parameters (tied input/output embeddings); "32m" in the repo name is the recipe name ("Glint 32M r6"), not the parameter count. Architecture: 15 layers, d_model 384, 6 heads, context 1024, Muon optimizer, batch 32, seed 1337, trainer train_gpt_ref_r6.py (0223c083).
Both arms saw the same ~12.4 GB of Polish text (public dataset SlayerLab/slayer-pl-8x3b @ d50df04a, files pack-00, pack-01 and shared/{core,core_sa,legal,legal_sa}; validation = pack-07 after 13-gram decontamination); B splits it more finely (on the training text 3.96 vs 4.25 bytes per token incl. EOS; on validation 4.14 vs 4.41), so it needed more steps for the same bytes. For the same reason B has lower bits per token (val 4.395 vs 4.668), which says nothing about model quality.
How to compare: loss per token is not comparable across tokenizers. We compare bits per byte on the same bytes (lower is better):
| A | B | B − A | |
|---|---|---|---|
| validation (7 232 820 B) | 1.0583 | 1.0622 | +0.0040 (+0.37 %) |
| held-out W (1 430 217 B) | 1.0781 | 1.0883 | +0.0101 (+0.94 %) |
From ~4 B bytes onward the gap is roughly constant (val +0.003…+0.004, W +0.008…+0.010) and does not close by the end of training. Curve every 10k steps (bits per byte; note: at the same step B has seen fewer bytes than A):
| step | A val | B val | A W | B W |
|---|---|---|---|---|
| 10 000 | 1.1910 | 1.2009 | 1.2327 | 1.2463 |
| 20 000 | 1.1448 | 1.1523 | 1.1799 | 1.1896 |
| 30 000 | 1.1200 | 1.1278 | 1.1499 | 1.1634 |
| 40 000 | 1.1032 | 1.1109 | 1.1304 | 1.1435 |
| 50 000 | 1.0888 | 1.0969 | 1.1141 | 1.1291 |
| 60 000 | 1.0773 | 1.0850 | 1.1000 | 1.1145 |
| 70 000 | 1.0675 | 1.0752 | 1.0885 | 1.1025 |
| 80 000 | 1.0610 | 1.0677 | 1.0818 | 1.0944 |
| 89 000 | 1.0583 | — | 1.0781 | — |
| 90 000 | — | 1.0633 | — | 1.0892 |
| 95 000 | — | 1.0622 | — | 1.0883 |
The same numbers, with checkpoint and data hashes, are in bpb_curve.json.
Caveat: one seed per arm, no noise estimate. The honest reading is "at this scale the B tokenizer brings no gain in bits per byte", not "B is worse".
Training data
Public dataset SlayerLab/slayer-pl-8x3b at revision d50df04a; the same files for both arms:
| file | use | sha256 |
|---|---|---|
pack-00/web.parquet |
training (web) | 5cbb6dfa3a276a30db3c2034105ffffa75cc29b767456e7d75c305f9bd7ec562 |
pack-01/web.parquet |
training (web) | 802f4829f41ef83e23979d2460500907bcfd44278e300eccb5e39ac666a04c2c |
shared/core.parquet |
training (core) | ea0b4b901ace3780501b02a3972f014f77ce722315476e6f93681616958e0a99 |
shared/core_sa.parquet |
training (core, CC BY-SA) | 910b18eeebbc26efbbd76554d7484d016c6df692a3b6a18492b5c9ceb32549cb |
shared/legal.parquet |
training (legal) | 6d4cc2f604d76452cd7e7ea52d1a04826f5e44f09bb035ba0853232fe373e74a |
shared/legal_sa.parquet |
training (legal, CC BY-SA) | a1609cd31e12e0f6fdfdb030edb538bc230fc5f5ab376d23136180338e9dbd1e |
pack-07/web.parquet |
validation (13-gram decontaminated against training) | 3d93331061595949f2e2cab187cfe688a117785c048e94688808b399b3f53b67 |
Usage
The model is a custom PyTorch architecture (same code as SlayerLab/GoLLeM-v6-250M), not a transformers class.
| file | content |
|---|---|
model.safetensors |
final weights, step 95 000 (fp32, 38,866,958 parameters; output head stored as a copy of the tied embedding) |
config.json |
architecture |
tokenizer.json |
arm B tokenizer by stubbornGuy, 32 000 tokens (tokenizers) |
modeling_gollem_v6.py |
model code (inference) and a small generate helper |
step_10000/ … step_90000/ |
intermediate checkpoints every 10k steps (same three files each; weights only, no optimizer state) |
bpb_curve.json |
bits-per-byte curve (step, checkpoint sha256, val, W) |
SHA256SUMS |
sha256 of every file |
# pip install torch safetensors tokenizers huggingface_hub
from huggingface_hub import snapshot_download
import sys
path = snapshot_download("SlayerLab/TokenizerTest-32m-B") # pin revision="<commit>" for reproducibility
sys.path.insert(0, path)
from modeling_gollem_v6 import load_gollem_v6, generate
model, tok = load_gollem_v6(path) # final checkpoint
print(generate(model, tok, "Kraków jest miastem", max_new_tokens=40, temperature=0.8, top_k=50, seed=1))
model_10k, _ = load_gollem_v6(f"{path}/step_10000") # an intermediate checkpoint
Loading any model.safetensors of this repository into modeling_gollem_v6.py reproduces the logits of the
corresponding training checkpoint exactly (checked with torch.equal for every step).
Limitations
- Small base model (~38.9M parameters): it continues Polish text; it does not answer questions or follow instructions. Limited knowledge; may reproduce errors and biases of web text. Research use only.
- Single run, single seed per arm; see the caveat above.
License
Weights: CC BY-SA 4.0 (the training text includes CC BY-SA parts of SlayerLab/slayer-pl-8x3b). Attribution:
TokenizerTest-32m-B, Arkadiusz Słota / Fabryka AI, link to this repository; derivative weights under the same licence.
The tokenizer (tokenizer.json) was created by stubbornGuy and is released here under CC BY-SA 4.0 with his consent.
Training data keep their upstream licences; see the dataset card of SlayerLab/slayer-pl-8x3b.
TokenizerTest — Fabryka AI. Author: Arkadiusz Słota. Tokenizer of arm B: stubbornGuy.
- Downloads last month
- 232