Substrate Language Modeling: publication artifacts

This repository contains the large reproducibility artifacts for the STLM Stage-1 pilot reported in Toward Compulsory Scene-Mediated Language Modeling: Meaningful Substrates for Every Next Word.

The paper, source code, tests, and reproduction scripts are maintained at github.com/emergent-wisdom/substrate-language-modeling. Download this artifact repository into the root of a GitHub checkout to restore the recorded stlm_prototype/data/ and stlm_prototype/runs/ paths.

Contents

  • The complete 1,024-token canonical vocabulary: native and downsampled images, pixel tensors, description tensors, vocabulary metadata, and symbolic descriptions.
  • Exactly 11 checkpoints underlying the reported direct-token, pixel, description, stateless-readout, and causal-readout measurements.
  • The four reported run manifests, compact result files, per-example TR-C predictions, and authenticated release records.
  • SHA256SUMS, covering every uploaded scientific artifact.

Local environments, caches, logs, smoke tests, superseded diagnostic runs, and unreported loose checkpoints are intentionally excluded.

Evidence boundary

The pilot establishes that a compulsory substrate interface can carry next-token information. It does not demonstrate semantic use or persistent scene emergence. The causal and stateless readouts also differ in architecture, capacity, compute, and supervision density, so their observed difference is not a pure estimate of history or memory.

License

These released research artifacts are available under CC BY 4.0, except where third-party terms apply. See LICENSE-CONTENT for scope and attribution.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support