Instructions to use yava-code/Tessera-1B-Nano-Base with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use yava-code/Tessera-1B-Nano-Base with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="yava-code/Tessera-1B-Nano-Base")# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("yava-code/Tessera-1B-Nano-Base") model = AutoModelForCausalLM.from_pretrained("yava-code/Tessera-1B-Nano-Base", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use yava-code/Tessera-1B-Nano-Base with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "yava-code/Tessera-1B-Nano-Base" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "yava-code/Tessera-1B-Nano-Base", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/yava-code/Tessera-1B-Nano-Base
- SGLang
How to use yava-code/Tessera-1B-Nano-Base with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "yava-code/Tessera-1B-Nano-Base" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "yava-code/Tessera-1B-Nano-Base", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "yava-code/Tessera-1B-Nano-Base" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "yava-code/Tessera-1B-Nano-Base", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use yava-code/Tessera-1B-Nano-Base with Docker Model Runner:
docker model run hf.co/yava-code/Tessera-1B-Nano-Base
Tessera-1B-Nano-Base
The plain half of a controlled experiment by Paragon Intelligence Labs: the
untouched HuggingFaceTB/SmolLM2-360M backbone, continued-pretrained with no
concept path, no extra parameters, no tricks. This is the matched NTP-only baseline of the comparison - every difference from its sibling is the concept path's doing, and nothing else.
TL;DR
This arm is the boring half on purpose: same tokens, same order, same initialization as the concept arm, no restarts, no NaNs. Science needs a control before it needs a result.
- Read this card as the yardstick: the final numbers below are what the concept arm is measured against.
- The comparison outcome and the intervention readings live in the whitepaper at
docs/whitepaper.mdin the ncp-smol repository.
Numbers
The raw record this narrative is built from; per-arm READMEs and the whitepaper hold the full analysis.
| Metric | Value |
|---|---|
| Held-out NTP loss | 2.5135 |
| Held-out perplexity | 12.3485 |
| Training tokens | 999,948,288 |
| Tracked compute estimate | $24.88 |
Intervention deltas are increases in held-out NTP loss relative to normal predicted concept feedback, evaluated on identical batches.
Architecture
- chunk size: 4
- product code: 15 segments x 64 entries
- causal concept blocks: 2
- injection point: before token decoder block 2
- NCP target: next continuous concept
- loss:
L_ntp + 1 L_ncp + 1 L_vq
This is a compact ConceptLM-style implementation, not an 8.9B NCP-ArchPreview replica. It omits iterative residual coding, cross-scale residual connections, and the large-scale training recipe.
Training data and provenance
The matched corpus and the training record are pinned in the ncp-smol repository:
packed-cache SHA256 hashes, the complete metric log, trainer state, and the raw
intervention evaluation JSON (also shipped in this repository as eval.json and
metrics.jsonl).
Loading
from transformers import AutoModelForCausalLM, AutoTokenizer
tokenizer = AutoTokenizer.from_pretrained("yava-code/Tessera-1B-Nano-Base")
model = AutoModelForCausalLM.from_pretrained("yava-code/Tessera-1B-Nano-Base")
The Tessera family
Three checkpoints, one story, in reading order:
- Tessera-135M-Gate - 135M architecture gate (overfit, published for the scale story)
- Tessera-1B-Nano-Base - 1B-token matched NTP-only control (this model)
- Tessera-1B-Nano - 1B-token matched concept arm
All cards are generated from the run artifacts by the same build_card; the
study repository holds the
whitepaper and full records.
References
- Code and study: https://github.com/yava-code/Tessera-1B-Nano
- ConceptLM: https://arxiv.org/abs/2602.08984
- NCP-ArchPreview: https://arxiv.org/abs/2609.10715
- Downloads last month
- 932