Tessera-135M-Gate

The smallest member of the Tessera family, and the reason the bigger ones exist. This 135M architecture-gate checkpoint (a HuggingFaceTB/SmolLM2-135M backbone with the full Next Concept Prediction path) was trained on a deliberately overfit TinyStories subset before any large spend, to prove the whole path - concept pooling, quantized codes, causal concept blocks, feedback - trains end to end. It is not a language-model quality result, and it is not meant to be: what makes it worth publishing is what broke, and what the break predicted at scale.

TL;DR

Small and cheap on purpose: this gate run exists to de-risk the 1B-token comparison before spending on it. It earned its keep in both directions.

  • All three objectives train (NTP, NCP, VQ fell 41 to 43% over the run), and the decoder causally uses the concept channel even at 135M: zeroing the feedback costs +0.0205 nats of held-out NTP loss.
  • It also caught a real failure mode: effective codebook perplexity stayed near 2.4 with usage around a third of the codebook, and shuffled feedback cost nothing - a low-entropy shortcut the 1B-token run later left far behind. Gate small, then scale: some behaviors only appear above a scale threshold.

Numbers

The raw record this narrative is built from; per-arm READMEs and the whitepaper hold the full analysis.

Metric Value
Held-out NTP loss 1.8751
Held-out perplexity 6.5213
Zero feedback delta 0.0205
Shuffled feedback delta 0.0000
Codebook perplexity / usage 2.4087 / 33.0%
Training tokens 16,777,216
Tracked compute estimate $1.38

Intervention deltas are increases in held-out NTP loss relative to normal predicted concept feedback, evaluated on identical batches.

Architecture

  • chunk size: 4
  • product code: 9 segments x 64 entries
  • causal concept blocks: 2
  • injection point: before token decoder block 2
  • NCP target: next continuous concept
  • loss: L_ntp + 1 L_ncp + 1 L_vq

This is a compact ConceptLM-style implementation, not an 8.9B NCP-ArchPreview replica. It omits iterative residual coding, cross-scale residual connections, and the large-scale training recipe.

Training data and provenance

The matched corpus and the training record are pinned in the ncp-smol repository: packed-cache SHA256 hashes, the complete metric log, trainer state, and the raw intervention evaluation JSON (also shipped in this repository as eval.json and metrics.jsonl).

Loading

from transformers import AutoModelForCausalLM, AutoTokenizer

tokenizer = AutoTokenizer.from_pretrained("yava-code/Tessera-135M-Gate")
model = AutoModelForCausalLM.from_pretrained("yava-code/Tessera-135M-Gate", trust_remote_code=True)

The Tessera family

Three checkpoints, one story, in reading order:

All cards are generated from the run artifacts by the same build_card; the study repository holds the whitepaper and full records.

References

Downloads last month
468
Safetensors
Model size
0.1B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for yava-code/Tessera-135M-Gate

Finetuned
(945)
this model

Collection including yava-code/Tessera-135M-Gate

Papers for yava-code/Tessera-135M-Gate