Substrate Language Modeling: publication artifacts
This repository contains the large reproducibility artifacts for the STLM Stage-1 pilot reported in Toward Compulsory Scene-Mediated Language Modeling: Meaningful Substrates for Every Next Word.
The paper, source code, tests, and reproduction scripts are maintained at
github.com/emergent-wisdom/substrate-language-modeling.
Download this artifact repository into the root of a GitHub checkout to restore
the recorded stlm_prototype/data/ and stlm_prototype/runs/ paths.
Contents
- The complete 1,024-token canonical vocabulary: native and downsampled images, pixel tensors, description tensors, vocabulary metadata, and symbolic descriptions.
- Exactly 11 checkpoints underlying the reported direct-token, pixel, description, stateless-readout, and causal-readout measurements.
- The four reported run manifests, compact result files, per-example TR-C predictions, and authenticated release records.
SHA256SUMS, covering every uploaded scientific artifact.
Local environments, caches, logs, smoke tests, superseded diagnostic runs, and unreported loose checkpoints are intentionally excluded.
Evidence boundary
The pilot establishes that a compulsory substrate interface can carry next-token information. It does not demonstrate semantic use or persistent scene emergence. The causal and stateless readouts also differ in architecture, capacity, compute, and supervision density, so their observed difference is not a pure estimate of history or memory.
License
These released research artifacts are available under CC BY 4.0, except where
third-party terms apply. See LICENSE-CONTENT for scope and attribution.