Goblin SLM 250M

Fresh Edge model with the SLM 32,768-token tokenizer. Target: 10 billion processed training targets. This is an experiment in progress, not a claim that 10B targets have completed or that the model is ready for chat.

The current segment is slm-10b-h200-v1, one H200 SXM Community Cloud at $3.59/hour. It resumes the verified A100 checkpoint at step 5,395 and 176,597,858 processed targets. Weights, optimizer, sampler and the full 10B schedule are preserved. The original $33.36 GPU budget is shared across machines; the H200 allocation has a $32.0075 ceiling including setup, plus storage, with shutdown requested for 28 September 2026 02:41 Amsterdam. The corpus identity, guarded optimizer and full-model restoration checks passed on the H200; training has resumed from that checkpoint. The earlier RTX 5090 Community and Secure allocation attempts were rejected for capacity. The first trained segment was slm-10b-a100-v3, on a new host after an input-boundary validation failure and unavailable restart capacity. It started from random weights and has now stopped with its final HF backup verified. Its checkpoints preserve optimizer state and progress on the full 10B schedule. Training uses the guarded ANVIL v3 FP32 matrix update, exact cross-entropy, BF16, compiled FlexAttention, MTP and fixed 4,096-token context.

Inputs are pinned public pretokenized shards: 50% FineWeb-Edu, 20% Python, 20% Cosmopedia, 10% math sampling weights. The selected pool is approximately 10.74B raw tokens before dropping incomplete edge documents and exact held-out duplicates. Random-window sampling is with replacement; processed targets are not distinct tokens seen. The source selection and tokenizer revision are saved under the run's launch and data folders.

Development and final evaluation use separate upstream shards; exact token documents from these splits are excluded from training. Framing uses adjacent EOS/BOS markers, preserving standalone marker IDs found literally in source text; literal adjacent pairs remain ambiguous. Near duplicates, repository overlap and benchmark contamination have not been fully audited. The final split is reserved until the complete 10B target. Data licenses and provenance remain those of the respective upstream sources; this repository does not redistribute the input shards or assert a single license for them.

This tokenizer is incompatible with earlier Goblin checkpoints. These weights start from random initialization; the previous Edge-2B model remains separate. This is a custom PyTorch checkpoint format, not a Transformers chat model. Checkpoint folders contain state.pt, complete.json and a hashed transfer manifest. Public checkpoints have verified remote bytes before being recorded as backed up. Run metrics and configuration are published separately.

No useful coding, reasoning or conversational capability has yet been established for this fresh run. Compare only matched evaluations; tokenizer changes prevent direct comparison of token loss with the earlier Goblin tokenizer.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Dataset used to train webstar34/goblin-slm-250m