Goblin SLM 250M
Fresh Edge model with the SLM 32,768-token tokenizer. Target: 10 billion processed training targets. This is an experiment in progress, not a claim that 10B targets have completed or that the model is ready for chat.
The current segment is slm-10b-h200-v1, one H200 SXM Community Cloud at
$3.59/hour. It resumes the verified A100 checkpoint at step 5,395 and
176,597,858 processed targets. Weights, optimizer, sampler and the full 10B
schedule are preserved. The original $33.36 GPU budget is shared across machines;
the H200 allocation has a $32.0075 ceiling including setup, plus storage, with
shutdown requested for 28 September 2026 02:41 Amsterdam. The corpus identity, guarded optimizer and full-model
restoration checks passed on the H200; training has resumed from that checkpoint.
The earlier RTX 5090 Community and Secure allocation attempts were rejected for
capacity. The first trained segment was slm-10b-a100-v3, on a new host after an
input-boundary validation failure and unavailable restart capacity. It started
from random weights and has now stopped with its final HF backup verified.
Its checkpoints preserve optimizer state and progress on the full 10B schedule.
Training uses the guarded ANVIL v3 FP32 matrix update, exact cross-entropy,
BF16, compiled FlexAttention, MTP and fixed 4,096-token context.
Inputs are pinned public pretokenized shards: 50% FineWeb-Edu, 20% Python,
20% Cosmopedia, 10% math sampling weights. The selected pool is approximately
10.74B raw tokens before dropping incomplete edge documents and exact held-out
duplicates. Random-window sampling is with replacement; processed targets are
not distinct tokens seen. The source selection and tokenizer revision are saved
under the run's launch and data folders.
Development and final evaluation use separate upstream shards; exact token documents from these splits are excluded from training. Framing uses adjacent EOS/BOS markers, preserving standalone marker IDs found literally in source text; literal adjacent pairs remain ambiguous. Near duplicates, repository overlap and benchmark contamination have not been fully audited. The final split is reserved until the complete 10B target. Data licenses and provenance remain those of the respective upstream sources; this repository does not redistribute the input shards or assert a single license for them.
This tokenizer is incompatible with earlier Goblin checkpoints. These weights
start from random initialization; the previous Edge-2B model remains separate.
This is a custom PyTorch checkpoint format, not a Transformers chat model.
Checkpoint folders contain state.pt, complete.json and a hashed transfer
manifest. Public checkpoints have verified remote bytes before being recorded
as backed up. Run metrics and configuration are published separately.
No useful coding, reasoning or conversational capability has yet been established for this fresh run. Compare only matched evaluations; tokenizer changes prevent direct comparison of token loss with the earlier Goblin tokenizer.