Tessera-1B-Nano-Base

The plain half of a controlled experiment by Paragon Intelligence Labs: the untouched HuggingFaceTB/SmolLM2-360M backbone, continued-pretrained with no concept path, no extra parameters, no tricks. This is the matched NTP-only baseline of the comparison - every difference from its sibling is the concept path's doing, and nothing else.

TL;DR

This arm is the boring half on purpose: same tokens, same order, same initialization as the concept arm, no restarts, no NaNs. Science needs a control before it needs a result.

  • Read this card as the yardstick: the final numbers below are what the concept arm is measured against.
  • The comparison outcome and the intervention readings live in the whitepaper at docs/whitepaper.md in the ncp-smol repository.

Numbers

The raw record this narrative is built from; per-arm READMEs and the whitepaper hold the full analysis.

Metric Value
Held-out NTP loss 2.5135
Held-out perplexity 12.3485
Training tokens 999,948,288
Tracked compute estimate $24.88

Intervention deltas are increases in held-out NTP loss relative to normal predicted concept feedback, evaluated on identical batches.

Architecture

  • chunk size: 4
  • product code: 15 segments x 64 entries
  • causal concept blocks: 2
  • injection point: before token decoder block 2
  • NCP target: next continuous concept
  • loss: L_ntp + 1 L_ncp + 1 L_vq

This is a compact ConceptLM-style implementation, not an 8.9B NCP-ArchPreview replica. It omits iterative residual coding, cross-scale residual connections, and the large-scale training recipe.

Training data and provenance

The matched corpus and the training record are pinned in the ncp-smol repository: packed-cache SHA256 hashes, the complete metric log, trainer state, and the raw intervention evaluation JSON (also shipped in this repository as eval.json and metrics.jsonl).

Loading

from transformers import AutoModelForCausalLM, AutoTokenizer

tokenizer = AutoTokenizer.from_pretrained("yava-code/Tessera-1B-Nano-Base")
model = AutoModelForCausalLM.from_pretrained("yava-code/Tessera-1B-Nano-Base")

The Tessera family

Three checkpoints, one story, in reading order:

All cards are generated from the run artifacts by the same build_card; the study repository holds the whitepaper and full records.

References

Downloads last month
932
Safetensors
Model size
0.4B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for yava-code/Tessera-1B-Nano-Base

Finetuned
(121)
this model
Quantizations
1 model

Space using yava-code/Tessera-1B-Nano-Base 1

Collection including yava-code/Tessera-1B-Nano-Base

Papers for yava-code/Tessera-1B-Nano-Base