veri-base-8739

1.2B params. Trained from scratch. Not a finetune, not a distill. These weights never existed before this run.

Training logs, benchmarks, architecture explorer: veriepic.dev/model

training

14B tokens, one pass each. FineWeb + DCLM. Cosine schedule that expired halfway (lesson learned: schedules that expire are banned), flat ever since.

The loss curve has a story. It collapsed to ~2.1 and flatlined for thousands of steps while samples stayed word salad β€” attention was broken from step 1 and the metric was lying. Fixed it, loss jumped back to ~7.5, real descent to 5.13 over 8739 steps. The shaded part is the lie. Everything right of the line is real.

architecture

24 identical blocks, dim 2048. RMSNorm, attention, residual, RMSNorm, SwiGLU, residual. Embeddings tied to the output head. RoPE positions, bf16, 4k context.

remy

Split 16 heads in half, full softmax on each half over the same input, subtract the second from the first with a learned per-pair lambda. Shared junk cancels, signal stays. QK-norm before RoPE, sliding window on even layers, 4 KV heads shared across 16 queries.

Shipped in the 1.19B: the remy family took 4 of the top 5 in the fixed 35-arch zoo, with qk-post edging first by 0.01. The big run committed to remy before the rerun finished β€” qk-post is data for the next scale, not a recall of this one.

terry

Trained with the private Terry build. Public Terry v2 lives in the main repo and matches it loss-for-loss: factored adaptive tables, grafted matrices, table nesterov, one momentum state per parameter.

duel terry muon adamw private
7M Β· 1000 steps 2.93 3.24 3.53 3.86 vs 3.84*
140M Β· H100 Β· 400 steps 3.95 4.09 4.59 3.95
states 7M 28.4MB 79.7MB 56.3MB 4.0MB
states 140M 226.9MB 380.4MB 452MB 61MB

*matched 300-step run. βˆ’9.5% vs muon, βˆ’16.9% vs adamw at 7M.

timmy

Custom 65k byte-level BPE, trained in 20 minutes on a rented CPU. Single digits split for arithmetic, think tags atomic.

test timmy 64k o200k cl100k
plain english 14 16 16
veri chat 23 29 29
reasoning 22 29 29
code 61 56 56
math 38 27 27
total 171 168 170

files

file size use
gguf/veri-base-8739-F16.gguf 2.4GB GGUF container, custom arch β€” loads with load_gguf_veri below, not stock runtimes
transformer/veri.bin 2.4GB torch / transformers
checkpoint/veri-base-8739.pt 3GB full bundle to keep training it
config.json, tokenizer.json β€” the usual
modeling_veri.py β€” the whole arch in one file, torch only
modeling_remy.py β€” standalone differential attention
# gguf container (custom arch β€” stock llama runtimes only know stock arches)
from modeling_veri import load_gguf_veri
model = load_gguf_veri("gguf/veri-base-8739-F16.gguf", dtype="bfloat16")
# transformers
from transformers import AutoModelForCausalLM, AutoTokenizer
tok = AutoTokenizer.from_pretrained("veriepic/Veri-Base", trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained("veriepic/Veri-Base", trust_remote_code=True)

Sampler settings (what the training log uses β€” greedy and penalty-free sampling babbles, this doesn't): temp 0.7, top-50, frequency penalty 0.2 applied as logits - 0.2 * counts before temperature.

ids = tok("Scientists have discovered that", return_tensors="pt").input_ids.cuda()
with torch.no_grad():
    for _ in range(32):
        lg = model(ids[:, -4096:]).logits[:, -1, :].float()
        counts = torch.bincount(ids[0], minlength=lg.size(-1)).float()
        lg = (lg - 0.2 * counts) / 0.7
        v, _ = torch.topk(lg, 50)
        lg = torch.where(lg < v[:, -1:], float("-inf"), lg)
        ids = torch.cat([ids, torch.softmax(lg, dim=-1).multinomial(1)], dim=1)
print(tok.decode(ids[0]))

status

Base weights. Pretraining parked (shards exhausted), SFT in progress. Fluent English, shaky facts β€” no instruction tuning yet.

veriepic.dev/model

Apache 2.0

Downloads last month
136
GGUF
Model size
1B params
Architecture
veri
Hardware compatibility
Log In to add your hardware

16-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Collection including veriepic/Veri-Base