TokenizerTest-32m-A (Polish research base model, tokenizer comparison — arm A)

A small Polish base language model trained as arm A of a controlled tokenizer experiment. It is a research checkpoint for comparing tokenizers, not a model for use. It is a Fabryka AI project (formerly SlayerLab).

Author: Arkadiusz Słota (Fabryka AI), who designed and ran the experiment. Tokenizer of arm B: stubbornGuy.

Companion model (tokenizer A/B experiment)

This model is one half of a controlled tokenizer comparison. Two models were trained with the same architecture, training recipe, seed and the same Polish text. The only difference is the tokenizer:

arm tokenizer repo final step
A Fabryka AI BPE, vocab 32 000 (60e23148) SlayerLab/TokenizerTest-32m-A 89 000
B tokenizer by stubbornGuy, vocab 32 000 (77eeebed) SlayerLab/TokenizerTest-32m-B 95 000

This model: A.

Shared recipe: ~38.9M parameters (tied input/output embeddings); "32m" in the repo name is the recipe name ("Glint 32M r6"), not the parameter count. Architecture: 15 layers, d_model 384, 6 heads, context 1024, Muon optimizer, batch 32, seed 1337, trainer train_gpt_ref_r6.py (0223c083).

Both arms saw the same ~12.4 GB of Polish text (public dataset SlayerLab/slayer-pl-8x3b @ d50df04a, files pack-00, pack-01 and shared/{core,core_sa,legal,legal_sa}; validation = pack-07 after 13-gram decontamination); B splits it more finely (on the training text 3.96 vs 4.25 bytes per token incl. EOS; on validation 4.14 vs 4.41), so it needed more steps for the same bytes. For the same reason B has lower bits per token (val 4.395 vs 4.668), which says nothing about model quality.

How to compare: loss per token is not comparable across tokenizers. We compare bits per byte on the same bytes (lower is better):

A B B − A
validation (7 232 820 B) 1.0583 1.0622 +0.0040 (+0.37 %)
held-out W (1 430 217 B) 1.0781 1.0883 +0.0101 (+0.94 %)

From ~4 B bytes onward the gap is roughly constant (val +0.003…+0.004, W +0.008…+0.010) and does not close by the end of training. Curve every 10k steps (bits per byte; note: at the same step B has seen fewer bytes than A):

step A val B val A W B W
10 000 1.1910 1.2009 1.2327 1.2463
20 000 1.1448 1.1523 1.1799 1.1896
30 000 1.1200 1.1278 1.1499 1.1634
40 000 1.1032 1.1109 1.1304 1.1435
50 000 1.0888 1.0969 1.1141 1.1291
60 000 1.0773 1.0850 1.1000 1.1145
70 000 1.0675 1.0752 1.0885 1.1025
80 000 1.0610 1.0677 1.0818 1.0944
89 000 1.0583 — 1.0781 —
90 000 — 1.0633 — 1.0892
95 000 — 1.0622 — 1.0883

The same numbers, with checkpoint and data hashes, are in bpb_curve.json.

Caveat: one seed per arm, no noise estimate. The honest reading is "at this scale the B tokenizer brings no gain in bits per byte", not "B is worse".

Training data

Public dataset SlayerLab/slayer-pl-8x3b at revision d50df04a; the same files for both arms:

file use sha256
pack-00/web.parquet training (web) 5cbb6dfa3a276a30db3c2034105ffffa75cc29b767456e7d75c305f9bd7ec562
pack-01/web.parquet training (web) 802f4829f41ef83e23979d2460500907bcfd44278e300eccb5e39ac666a04c2c
shared/core.parquet training (core) ea0b4b901ace3780501b02a3972f014f77ce722315476e6f93681616958e0a99
shared/core_sa.parquet training (core, CC BY-SA) 910b18eeebbc26efbbd76554d7484d016c6df692a3b6a18492b5c9ceb32549cb
shared/legal.parquet training (legal) 6d4cc2f604d76452cd7e7ea52d1a04826f5e44f09bb035ba0853232fe373e74a
shared/legal_sa.parquet training (legal, CC BY-SA) a1609cd31e12e0f6fdfdb030edb538bc230fc5f5ab376d23136180338e9dbd1e
pack-07/web.parquet validation (13-gram decontaminated against training) 3d93331061595949f2e2cab187cfe688a117785c048e94688808b399b3f53b67

Usage

The model is a custom PyTorch architecture (same code as SlayerLab/GoLLeM-v6-250M), not a transformers class.

file content
model.safetensors final weights, step 89 000 (fp32, 38,866,958 parameters; output head stored as a copy of the tied embedding)
config.json architecture
tokenizer.json arm A tokenizer, 32 000 tokens (tokenizers)
modeling_gollem_v6.py model code (inference) and a small generate helper
step_10000/ … step_80000/ intermediate checkpoints every 10k steps (same three files each; weights only, no optimizer state)
bpb_curve.json bits-per-byte curve (step, checkpoint sha256, val, W)
SHA256SUMS sha256 of every file
# pip install torch safetensors tokenizers huggingface_hub
from huggingface_hub import snapshot_download
import sys

path = snapshot_download("SlayerLab/TokenizerTest-32m-A")   # pin revision="<commit>" for reproducibility
sys.path.insert(0, path)
from modeling_gollem_v6 import load_gollem_v6, generate

model, tok = load_gollem_v6(path)                          # final checkpoint
print(generate(model, tok, "Kraków jest miastem", max_new_tokens=40, temperature=0.8, top_k=50, seed=1))
model_10k, _ = load_gollem_v6(f"{path}/step_10000")        # an intermediate checkpoint

Loading any model.safetensors of this repository into modeling_gollem_v6.py reproduces the logits of the corresponding training checkpoint exactly (checked with torch.equal for every step).

Limitations

  • Small base model (~38.9M parameters): it continues Polish text; it does not answer questions or follow instructions. Limited knowledge; may reproduce errors and biases of web text. Research use only.
  • Single run, single seed per arm; see the caveat above.

License

Weights: CC BY-SA 4.0 (the training text includes CC BY-SA parts of SlayerLab/slayer-pl-8x3b). Attribution: TokenizerTest-32m-A, Arkadiusz Słota / Fabryka AI, link to this repository; derivative weights under the same licence. Training data keep their upstream licences; see the dataset card of SlayerLab/slayer-pl-8x3b.

TokenizerTest — Fabryka AI. Author: Arkadiusz Słota.

Downloads last month
224
Safetensors
Model size
51.2M params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support