Zarya

A 50,852,096-parameter English base model from the Tiny llm ablation experiment. Trained from random initialization for 15,000 optimizer steps, processing 3,932,160,000 source tokens. These are processed input tokens, not unique text or supervised target-token counts.

Architecture

Zarya-inspired hybrid, trained from scratch: 9 layers, width 512, FFN width 1440, 8 query / 4 KV heads with head dimension 128, RoPE theta 1,000,000, QK normalization, RMSNorm epsilon 1e-6, tied embeddings and no Q50M residual gates. Vocabulary 32,769 includes a trained mask token. The attention projection width is 1024 despite hidden width 512. Benchmarks use clean causal AR scoring with the full vocabulary, including mask-token probability.

Tokenizer foundation: q-project/Q-50M-Base. This model is trained from scratch, not fine-tuned from the reference weights. Exact trained configuration and custom model code are included.

Training

  • Data: FineWeb-Edu, sample-10BT, local Parquet shards; shuffle buffer 100,000.
  • Objective: Slotted AR/masked-denoising hybrid; slot sizes [2, 4, 8, 16, 32, 64] with training-fraction boundaries [0.06, 0.2, 0.4, 0.6, 0.8, 1.0]; diffusion loss proportion 0.5. Visible and masked slots are shuffled separately, preserving original position IDs; attention is causal in permuted order. No time conditioning. Hybrid training loss is not autoregressive NLL.
  • Final recorded batch: 8 sequences × 16 gradient accumulation × 2048 tokens = 262,144 source tokens per update.
  • Fused AdamW, peak LR 0.001, betas (0.9, 0.95), weight decay 0.1 excluding bias/norm/1D parameters, gradient clipping 1.0. Warmup 150 steps, cosine decay to 10% of peak LR.
  • BF16 compute, one RTX 5070 Ti, seed 2026; released checkpoint weights remain FP32. Training configuration.

Evaluation

Full selected splits, lm-eval 0.4.12, no added few-shot examples or chat template, BF16 on RTX 5070 Ti, context cap 2048 (ArithMark 1024). New evaluations explicitly disable TF32; historical core manifests predate the explicit TF32 flag. Scores below are percentages; ± is one standard error, and the separate interval column is 95% CI where available.

Dataset Split Examples Metric Score ± SE (%) 95% CI (%)
HellaSwag validation 10,042 acc_norm 27.33 ± 0.44 [26.46, 28.21]
ARC-Easy test 2,376 acc_norm 39.94 ± 1.01 [37.99, 41.93]
ARC-Challenge test 1,172 acc_norm 23.46 ± 1.24 [21.13, 25.97]
PIQA validation 1,838 acc_norm 57.73 ± 1.15 [55.45, 59.97]
WinoGrande validation 1,267 acc 52.49 ± 1.40 [49.73, 55.22]
OpenBookQA test 500 acc_norm 26.60 ± 1.98 [22.92, 30.64]
BoolQ validation 3,270 acc 59.24 ± 0.86 [57.54, 60.91]
LAMBADA OpenAI test 5,153 acc 18.46 ± 0.54 [17.42, 19.54]
ArithMark-3 train 1,000 acc_norm 36.00 ± 1.52 [33.08, 39.02]
Balanced COPA train 1,000 acc 51.90 ± 1.58 [48.80, 54.98]
BLiMP train 67,000 acc 71.97 ± 0.15 —
CommonsenseQA validation 1,221 acc 19.66 ± 1.14 [17.52, 21.98]
MMLU continuation test 14,042 acc 25.00 ± 0.36 —
SciQ (with support) test 1,000 acc_norm 62.40 ± 1.53 [59.36, 65.35]
TruthfulQA MC2 validation 817 acc 45.86 ± 1.55 —
BananaMind Base 1.1 test 350 raw_accuracy 44.00 ± 2.66 [38.89, 49.24]

LAMBADA requires all final-word tokens to match. The descriptive macro-average is 41.3754%: one primary metric per each of the 16 tasks, with equal task weight. PPL, standard errors, CI bounds and MMLU/BLiMP leaf scores are excluded. It is not an official leaderboard score or a statistical ranking test.

Full selected splits; lm-eval 0.4.12; seed 1234; BF16 on RTX 5070 Ti; context cap 2048 (ArithMark 1024), TF32 disabled on new runs (historical core flag unrecorded), no chat template and no added few-shot examples. TruthfulQA retains the harness's fixed six-QA preamble. ArithMark and BananaMind normalize by continuation token count; ordinary harness acc_norm uses its own length normalization. BananaMind is raw accuracy, not Elo. SciQ includes the support passage. Balanced COPA uses the mirrored 1000-item train-named evaluation split; cRia's split was inferred, not confirmed. MMLU scores full answer continuations across 57 subjects, weighted by item count; BLiMP averages 67 equal-sized minimal-pair subsets. Standard errors are retained from each evaluator. Wilson intervals are reported only where the runner logged binary item accuracy; MC2 is probability mass, not binary accuracy. These intervals do not model dependence between paired/templated examples or training seed variation. UL2, PrefixLM and experimental diffusion PLL use their documented conditional scoring protocols; PLL exposes the other answer tokens and is not autoregressive likelihood. cRia's published scores used a different precision and benchmark-adapted checkpoint; this completes our comparison coverage, not an independent reproduction of cRia or an official leaderboard submission.

WikiText-2 raw test continuation, reported separately: 291 nonoverlapping blocks of 1024 tokens, 512-token prefix plus 512-token scored suffix; 148,992 scored tokens, 335 tail tokens excluded. GPU evaluation; percentile block bootstrap with 10,000 resamples and seed 2026. PPL intervals exponentiate NLL endpoints. This is not rolling or word perplexity.

  • torch.bfloat16: NLL 3.830518, 95% CI [3.789140, 3.871400]; PPL 46.086, 95% CI [44.218, 48.010].
  • torch.float32: NLL 3.830018, 95% CI [3.788650, 3.870869]; PPL 46.063, 95% CI [44.197, 47.984].

Structured results and provenance, raw task outputs and per-block continuation scores. HF metadata contains author-reported measurements; no verified benchmark badge is claimed. Historical and new measurements refer to identical weight hashes; individual run manifests retain their original dates and source hashes.

Usage

Install requirements.txt; tested with PyTorch 2.11.0 and Transformers 5.17.0. Loading custom code requires trust_remote_code=True. Example on CPU:

from transformers import AutoTokenizer, AutoModelForCausalLM
repo = "d0rj/zarya-51M-base"
tokenizer = AutoTokenizer.from_pretrained(repo)
model = AutoModelForCausalLM.from_pretrained(repo, trust_remote_code=True).eval()
inputs = tokenizer("The capital of France is", return_tensors="pt")
output = model.generate(**inputs, max_new_tokens=32, do_sample=False)
print(tokenizer.decode(output[0], skip_special_tokens=True))

After downloading this repository, install the evaluation dependencies and authenticate for the gated BananaMind dataset after accepting its terms:

pip install -r evaluation/repro/requirements.txt
hf auth login
python evaluation/repro/run.py --device cuda:0 --dtype bfloat16 --batch-size 8 --output comparison-rerun
python evaluation/repro/continuation.py --device cuda:0 --dtype bfloat16 --output continuation-bf16.json
python evaluation/repro/continuation.py --device cuda:0 --dtype float32 --output continuation-fp32.json

--limit / --limit-blocks are smoke checks only. Dataset examples are not redistributed. The bundled runners use this published model and tokenizer.

Training logs and limitations

TensorBoard event files contain available training telemetry and the full evaluation at optimizer step 15,000. train/loss has 750 visible points spanning steps 20–15,000. Resume/purge records are retained. Comet experiment. Evaluation exports retain numeric metrics, uncertainty and task/subtask breakdowns.

Equal source-token budgets do not imply equal target supervision, parameter count or FLOPs. Training dataset file revisions and overlap with benchmarks were not independently audited. Results come from one training seed; uncertainty intervals do not capture training-seed variability or dependence between templated examples. These are small base models, not instruction-tuned assistants.

Downloads last month
98
Safetensors
Model size
50.9M params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Dataset used to train d0rj/zarya-51M-base

Collection including d0rj/zarya-51M-base

Evaluation results