Instructions to use yava-code/Tessera-135M-Gate with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use yava-code/Tessera-135M-Gate with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="yava-code/Tessera-135M-Gate", trust_remote_code=True)# Load model directly from transformers import AutoModelForCausalLM model = AutoModelForCausalLM.from_pretrained("yava-code/Tessera-135M-Gate", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use yava-code/Tessera-135M-Gate with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "yava-code/Tessera-135M-Gate" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "yava-code/Tessera-135M-Gate", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/yava-code/Tessera-135M-Gate
- SGLang
How to use yava-code/Tessera-135M-Gate with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "yava-code/Tessera-135M-Gate" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "yava-code/Tessera-135M-Gate", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "yava-code/Tessera-135M-Gate" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "yava-code/Tessera-135M-Gate", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use yava-code/Tessera-135M-Gate with Docker Model Runner:
docker model run hf.co/yava-code/Tessera-135M-Gate
Tessera-135M-Gate
The smallest member of the Tessera family, and the reason the bigger ones exist.
This 135M architecture-gate checkpoint (a HuggingFaceTB/SmolLM2-135M backbone
with the full Next Concept Prediction path) was trained on a deliberately overfit
TinyStories subset before any large spend, to prove the whole path - concept
pooling, quantized codes, causal concept blocks, feedback - trains end to end.
It is not a language-model quality result, and it is not meant to be: what makes
it worth publishing is what broke, and what the break predicted at scale.
TL;DR
Small and cheap on purpose: this gate run exists to de-risk the 1B-token comparison before spending on it. It earned its keep in both directions.
- All three objectives train (NTP, NCP, VQ fell 41 to 43% over the run), and the decoder causally uses the concept channel even at 135M: zeroing the feedback costs +0.0205 nats of held-out NTP loss.
- It also caught a real failure mode: effective codebook perplexity stayed near 2.4 with usage around a third of the codebook, and shuffled feedback cost nothing - a low-entropy shortcut the 1B-token run later left far behind. Gate small, then scale: some behaviors only appear above a scale threshold.
Numbers
The raw record this narrative is built from; per-arm READMEs and the whitepaper hold the full analysis.
| Metric | Value |
|---|---|
| Held-out NTP loss | 1.8751 |
| Held-out perplexity | 6.5213 |
| Zero feedback delta | 0.0205 |
| Shuffled feedback delta | 0.0000 |
| Codebook perplexity / usage | 2.4087 / 33.0% |
| Training tokens | 16,777,216 |
| Tracked compute estimate | $1.38 |
Intervention deltas are increases in held-out NTP loss relative to normal predicted concept feedback, evaluated on identical batches.
Architecture
- chunk size: 4
- product code: 9 segments x 64 entries
- causal concept blocks: 2
- injection point: before token decoder block 2
- NCP target: next continuous concept
- loss:
L_ntp + 1 L_ncp + 1 L_vq
This is a compact ConceptLM-style implementation, not an 8.9B NCP-ArchPreview replica. It omits iterative residual coding, cross-scale residual connections, and the large-scale training recipe.
Training data and provenance
The matched corpus and the training record are pinned in the ncp-smol repository:
packed-cache SHA256 hashes, the complete metric log, trainer state, and the raw
intervention evaluation JSON (also shipped in this repository as eval.json and
metrics.jsonl).
Loading
from transformers import AutoModelForCausalLM, AutoTokenizer
tokenizer = AutoTokenizer.from_pretrained("yava-code/Tessera-135M-Gate")
model = AutoModelForCausalLM.from_pretrained("yava-code/Tessera-135M-Gate", trust_remote_code=True)
The Tessera family
Three checkpoints, one story, in reading order:
- Tessera-135M-Gate - 135M architecture gate (overfit, published for the scale story) (this model)
- Tessera-1B-Nano-Base - 1B-token matched NTP-only control
- Tessera-1B-Nano - 1B-token matched concept arm
All cards are generated from the run artifacts by the same build_card; the
study repository holds the
whitepaper and full records.
References
- Code and study: https://github.com/yava-code/Tessera-1B-Nano
- ConceptLM: https://arxiv.org/abs/2602.08984
- NCP-ArchPreview: https://arxiv.org/abs/2609.10715
- Downloads last month
- 468
Model tree for yava-code/Tessera-135M-Gate
Base model
HuggingFaceTB/SmolLM2-135M