Tercet-R-1.0

Reasoning + tool-call chat model (~502M) — Tercet-R family

Model Stage License Base

A ~502M hybrid GDN-2 + GQA model, supervised fine-tuned for thinking and tools


What this is

Tercet-R-1.0 is the first public reasoning / tool-use checkpoint in the Tercet-R line.


Chat contract

Thinking

Each assistant turn is prefixed with a zero-loss control token:

Mode Prefix Typical body
think <|think|>\n <think>…</think> then the answer
no-think <|no_think|>\n answer only

inference.py and the local scripts/chat.py stream the <think> region live (dim yellow) and hide the control tokens.

Tool calls (SmolTalk JSON)

<tool_call>
{"name": "web_search", "arguments": {"query": "…"}}
</tool_call>

Tool results

Each observation is a tool (or user) turn prefixed with the special token:

<|tool_response|>
{observation}

inference.py --tools web_search pauses after a <tool_call>, you paste the search result, and generation continues.


Install & run

pip install torch safetensors tokenizers huggingface_hub
hf download kerzgrr/Tercet-R-1.0 inference.py --local-dir .
python inference.py
python inference.py --prompt "What is the capital of France?"
python inference.py --tools web_search
python inference.py --no-think --prompt "Reply in one sentence."

inference.py auto-downloads weights / tokenizer / tiny_gdn/ and auto-installs pinned flash-linear-attention. Git is required on PATH.

Flag Default Description
--prompt One-shot user message
--system System prompt, used verbatim
--think / --no-think think Assistant control prefix
--tools Built-in tools (web_search)
--temperature 0.7 Sampling temperature
--max-new-tokens 4096 Max generation length
--device cuda if available cuda / cpu

Interactive commands: /think /no_think /system … /reset /exit.


Model architecture

Same TinyGDN hybrid as Tercet-base (501,635,264 parameters):

Layers 32 (GDN-2 ×3 + GQA every 4th)
Hidden 1,024
MLP SwiGLU 2,624
Attention 8 Q / 2 KV, head dim 128, partial RoPE
Linear Gated DeltaNet-2, 8 heads × 128
Vocab 49,152 BPE
Context 16,384

Training

Stage Details
Base HuggingFaceFW/fineweb-eduTercet-base
Mid-SFT HuggingFaceTB/smoltalk2 Mid: Llama-Nemotron-Post-Training-Dataset + OpenThoughts3-1.2M
Instruct SFT HuggingFaceTB/smoltalk2 SFT (SmolTalk, OpenHermes-2.5, OpenThoughts3, Aya, Hermes function calling, s1K, Tulu-3 personas IF, xLAM, LongAlign, Mixture-of-Thoughts, …), seq 16,384, AdamW 5×10⁻⁵, 61.1 hours
Checkpoint optimizer step 4,500 (latest complete instruct snapshot)
Weights EMA (this repo's model.safetensors)
Val loss (EMA) 1.744 (ppl 5.72)

Limitations

  • Scale: ~502M is a research / edge model, not a frontier system
  • Requires flash-linear-attention; not GGUF / llama.cpp compatible today

Model family

Model Stage Hub
Tercet-base Pretrain kerzgrr/Tercet-base
Tercet SFT chat kerzgrr/Tercet
Tercet-R-1.0 SFT reasoning + tools this repo

Citation

@misc{tercetr2026,
  title={Tercet-R-1.0: A 502M Hybrid GDN-2 + GQA Reasoning Model},
  author={kerzgrr},
  year={2026},
  url={https://huggingface.co/kerzgrr/Tercet-R-1.0}
}

R is for reasoning.

Downloads last month
579
Safetensors
Model size
0.5B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for kerzgrr/Tercet-R-1.0

Finetuned
(2)
this model

Datasets used to train kerzgrr/Tercet-R-1.0

Collection including kerzgrr/Tercet-R-1.0