Q-MoD

A 50,919,168-parameter English base model from the Tiny llm ablation experiment. Trained from random initialization for 15,000 optimizer steps, processing 3,932,160,000 source tokens. These are processed input tokens, not unique text or supervised target-token counts.

Architecture

10 decoder layers, width 512, 8 query / 2 KV heads, head dimension 64, RoPE, QK normalization, gated residuals and tied 32,768-token embeddings. Five alternating whole blocks (zero-based layers 1, 3, 5, 7, 9) use Mixture-of-Depths token routing. Training selects the top 12.5% of valid tokens per sequence; evaluation selects each token causally with router logit > 0, without future-dependent top-k. Routed FFN width 1797 (ordinary blocks 1792) makes the total exactly 50,919,168 parameters. This routes tokens through blocks, not experts. Inference compute depends on selected tokens.

Tokenizer foundation: q-project/Q-50M-Base. This model is trained from scratch, not fine-tuned from the reference weights. Exact trained configuration and custom model code are included.

Training

  • Data: FineWeb-Edu, sample-10BT, local Parquet shards; shuffle buffer 100,000.
  • Objective: Causal next-token cross-entropy plus router binary classification auxiliary loss, coefficient 0.01.
  • Final recorded batch: 4 sequences × 32 gradient accumulation × 2048 tokens = 262,144 source tokens per update.
  • Fused AdamW, peak LR 0.001, betas (0.9, 0.95), weight decay 0.1 excluding bias/norm/1D parameters, gradient clipping 1.0. Warmup 150 steps, cosine decay to 10% of peak LR.
  • BF16 compute, one RTX 5070 Ti, seed 2026; released checkpoint weights remain FP32. Training configuration.

Evaluation

Full selected splits, lm-eval 0.4.12, no added few-shot examples or chat template, BF16 on RTX 5070 Ti, context cap 2048 (ArithMark 1024). New evaluations explicitly disable TF32; historical core manifests predate the explicit TF32 flag. Scores below are percentages; ± is one standard error, and the separate interval column is 95% CI where available.

Dataset Split Examples Metric Score ± SE (%) 95% CI (%)
HellaSwag validation 10,042 acc_norm 28.75 ± 0.45 [27.87, 29.64]
ARC-Easy test 2,376 acc_norm 41.92 ± 1.01 [39.95, 43.91]
ARC-Challenge test 1,172 acc_norm 23.12 ± 1.23 [20.80, 25.62]
PIQA validation 1,838 acc_norm 58.76 ± 1.15 [56.49, 60.99]
WinoGrande validation 1,267 acc 51.38 ± 1.40 [48.63, 54.12]
OpenBookQA test 500 acc_norm 30.60 ± 2.06 [26.72, 34.77]
BoolQ validation 3,270 acc 55.78 ± 0.87 [54.07, 57.47]
LAMBADA OpenAI test 5,153 acc 16.69 ± 0.52 [15.70, 17.73]
ArithMark-3 train 1,000 acc_norm 34.10 ± 1.50 [31.23, 37.09]
Balanced COPA train 1,000 acc 53.30 ± 1.58 [50.20, 56.37]
BLiMP train 67,000 acc 73.18 ± 0.15 —
CommonsenseQA validation 1,221 acc 19.49 ± 1.13 [17.37, 21.81]
MMLU continuation test 14,042 acc 25.22 ± 0.37 —
SciQ (with support) test 1,000 acc_norm 61.20 ± 1.54 [58.14, 64.17]
TruthfulQA MC2 validation 817 acc 44.18 ± 1.52 —
BananaMind Base 1.1 test 350 raw_accuracy 45.71 ± 2.67 [40.57, 50.95]

LAMBADA requires all final-word tokens to match. The descriptive macro-average is 41.4613%: one primary metric per each of the 16 tasks, with equal task weight. PPL, standard errors, CI bounds and MMLU/BLiMP leaf scores are excluded. It is not an official leaderboard score or a statistical ranking test.

Full selected splits; lm-eval 0.4.12; seed 1234; BF16 on RTX 5070 Ti; context cap 2048 (ArithMark 1024), TF32 disabled on new runs (historical core flag unrecorded), no chat template and no added few-shot examples. TruthfulQA retains the harness's fixed six-QA preamble. ArithMark and BananaMind normalize by continuation token count; ordinary harness acc_norm uses its own length normalization. BananaMind is raw accuracy, not Elo. SciQ includes the support passage. Balanced COPA uses the mirrored 1000-item train-named evaluation split; cRia's split was inferred, not confirmed. MMLU scores full answer continuations across 57 subjects, weighted by item count; BLiMP averages 67 equal-sized minimal-pair subsets. Standard errors are retained from each evaluator. Wilson intervals are reported only where the runner logged binary item accuracy; MC2 is probability mass, not binary accuracy. These intervals do not model dependence between paired/templated examples or training seed variation. UL2, PrefixLM and experimental diffusion PLL use their documented conditional scoring protocols; PLL exposes the other answer tokens and is not autoregressive likelihood. cRia's published scores used a different precision and benchmark-adapted checkpoint; this completes our comparison coverage, not an independent reproduction of cRia or an official leaderboard submission.

WikiText-2 raw test continuation, reported separately: 291 nonoverlapping blocks of 1024 tokens, 512-token prefix plus 512-token scored suffix; 148,992 scored tokens, 335 tail tokens excluded. GPU evaluation; percentile block bootstrap with 10,000 resamples and seed 2026. PPL intervals exponentiate NLL endpoints. This is not rolling or word perplexity.

  • torch.bfloat16: NLL 3.689238, 95% CI [3.649657, 3.728491]; PPL 40.014, 95% CI [38.461, 41.616].
  • torch.float32: NLL 3.688851, 95% CI [3.649183, 3.728063]; PPL 39.999, 95% CI [38.443, 41.598].

Structured results and provenance, raw task outputs and per-block continuation scores. HF metadata contains author-reported measurements; no verified benchmark badge is claimed. Historical and new measurements refer to identical weight hashes; individual run manifests retain their original dates and source hashes.

Usage

Install requirements.txt; tested with PyTorch 2.11.0 and Transformers 5.17.0. Loading custom code requires trust_remote_code=True. Example on CPU:

from transformers import AutoTokenizer, AutoModelForCausalLM
repo = "d0rj/q-mod-51M-base"
tokenizer = AutoTokenizer.from_pretrained(repo)
model = AutoModelForCausalLM.from_pretrained(repo, trust_remote_code=True).eval()
inputs = tokenizer("The capital of France is", return_tensors="pt")
output = model.generate(**inputs, max_new_tokens=32, do_sample=False)
print(tokenizer.decode(output[0], skip_special_tokens=True))

After downloading this repository, install the evaluation dependencies and authenticate for the gated BananaMind dataset after accepting its terms:

pip install -r evaluation/repro/requirements.txt
hf auth login
python evaluation/repro/run.py --device cuda:0 --dtype bfloat16 --batch-size 8 --output comparison-rerun
python evaluation/repro/continuation.py --device cuda:0 --dtype bfloat16 --output continuation-bf16.json
python evaluation/repro/continuation.py --device cuda:0 --dtype float32 --output continuation-fp32.json

--limit / --limit-blocks are smoke checks only. Dataset examples are not redistributed. The bundled runners use this published model and tokenizer.

Training logs and limitations

TensorBoard event files contain available training telemetry and the full evaluation at optimizer step 15,000. train/loss has 750 visible points spanning steps 20–15,000. Resume/purge records are retained. Comet experiment. Evaluation exports retain numeric metrics, uncertainty and task/subtask breakdowns.

Equal source-token budgets do not imply equal target supervision, parameter count or FLOPs. Training dataset file revisions and overlap with benchmarks were not independently audited. Results come from one training seed; uncertainty intervals do not capture training-seed variability or dependence between templated examples. These are small base models, not instruction-tuned assistants.

Downloads last month
90
Safetensors
Model size
50.9M params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Dataset used to train d0rj/q-mod-51M-base

Collection including d0rj/q-mod-51M-base

Evaluation results