Diffusion v1 (failed; PLL)

Failed ablation: token-frequency/context-collapse behavior. Retained for reproducibility and negative results. The repaired model is diffusion v2.

A 51,392,512-parameter English base model from the Tiny llm ablation experiment. Trained from random initialization for 15,000 optimizer steps, processing 3,932,160,000 source tokens. These are processed input tokens, not unique text or supervised target-token counts.

Architecture

Failed v1 absorbing-mask diffusion ablation: 10 bidirectional layers, width 512, SwiGLU width 1536, 8 attention heads, head dimension 64, RoPE, RMSNorm and tied 32,768-token output vocabulary. A separate learned mask input vector uses ID 32768, outside the output vocabulary. Legacy additive time conditioning is retained exactly. The checkpoint stores 51,392,512 parameters.

Tokenizer foundation: q-project/Q-50M-Base. This model is trained from scratch, not fine-tuned from the reference weights. Exact trained configuration and custom model code are included.

Training

  • Data: FineWeb-Edu, sample-10BT, local Parquet shards; shuffle buffer 100,000.
  • Objective: Sample noise time uniformly; independently mask tokens with that probability, predict masked originals with inverse-time weighting normalized by source token count. This run collapsed toward token-frequency predictions and is published as a failed ablation. The repaired diffusion v2 disables additive time conditioning.
  • Final recorded batch: 8 sequences × 16 gradient accumulation × 2048 tokens = 262,144 source tokens per update.
  • Fused AdamW, peak LR 0.001, betas (0.9, 0.95), weight decay 0.1 excluding bias/norm/1D parameters, gradient clipping 1.0. Warmup 150 steps, cosine decay to 10% of peak LR.
  • BF16 compute, one RTX 5070 Ti, seed 2026; released checkpoint weights remain FP32. Training configuration.

Evaluation

Full selected splits, lm-eval 0.4.12, no added few-shot examples or chat template, BF16 on RTX 5070 Ti, context cap 2048 (ArithMark 1024). New evaluations explicitly disable TF32; historical core manifests predate the explicit TF32 flag. Scores below are percentages; ± is one standard error, and the separate interval column is 95% CI where available.

Dataset Split Examples Metric Score ± SE (%) 95% CI (%)
HellaSwag validation 10,042 pll_acc_norm 25.97 ± 0.44 [25.12, 26.84]
ARC-Easy test 2,376 pll_acc_norm 26.68 ± 0.91 [24.94, 28.50]
ARC-Challenge test 1,172 pll_acc_norm 24.06 ± 1.25 [21.70, 26.59]
PIQA validation 1,838 pll_acc_norm 50.05 ± 1.17 [47.77, 52.34]
WinoGrande validation 1,267 pll_acc 50.99 ± 1.40 [48.24, 53.73]
OpenBookQA test 500 pll_acc_norm 25.80 ± 1.96 [22.16, 29.81]
BoolQ validation 3,270 pll_acc 37.83 ± 0.85 [36.18, 39.50]
LAMBADA OpenAI test 5,153 pll_acc 0.00 ± 0.00 [0.00, 0.07]
ArithMark-3 train 1,000 pll_acc_norm 29.90 ± 1.45 [27.14, 32.81]
Balanced COPA train 1,000 pll_acc 50.00 ± 1.58 [46.91, 53.09]
BLiMP train 67,000 pll_acc 57.39 ± 0.16 —
CommonsenseQA validation 1,221 pll_acc 19.57 ± 1.14 [17.45, 21.89]
MMLU continuation test 14,042 pll_acc 23.38 ± 0.36 —
SciQ (with support) test 1,000 pll_acc_norm 23.70 ± 1.35 [21.17, 26.43]
TruthfulQA MC2 validation 817 pll_acc 48.99 ± 1.64 —
BananaMind Base 1.1 test 350 pll_raw_accuracy 26.29 ± 2.36 [21.95, 31.14]

These are experimental single-mask continuation PLL scores, not autoregressive likelihood. Other answer tokens remain visible. LAMBADA is whole-answer reconstruction under this protocol, not ordinary next-word generation. Compare this model in a separate diagnostic group. LAMBADA requires all final-word tokens to match. The descriptive macro-average is 32.5378%: one primary metric per each of the 16 tasks, with equal task weight. PPL, standard errors, CI bounds and MMLU/BLiMP leaf scores are excluded. It is not an official leaderboard score or a statistical ranking test.

Full selected splits; lm-eval 0.4.12; seed 1234; BF16 on RTX 5070 Ti; context cap 2048 (ArithMark 1024), TF32 disabled on new runs (historical core flag unrecorded), no chat template and no added few-shot examples. TruthfulQA retains the harness's fixed six-QA preamble. ArithMark and BananaMind normalize by continuation token count; ordinary harness acc_norm uses its own length normalization. BananaMind is raw accuracy, not Elo. SciQ includes the support passage. Balanced COPA uses the mirrored 1000-item train-named evaluation split; cRia's split was inferred, not confirmed. MMLU scores full answer continuations across 57 subjects, weighted by item count; BLiMP averages 67 equal-sized minimal-pair subsets. Standard errors are retained from each evaluator. Wilson intervals are reported only where the runner logged binary item accuracy; MC2 is probability mass, not binary accuracy. These intervals do not model dependence between paired/templated examples or training seed variation. UL2, PrefixLM and experimental diffusion PLL use their documented conditional scoring protocols; PLL exposes the other answer tokens and is not autoregressive likelihood. cRia's published scores used a different precision and benchmark-adapted checkpoint; this completes our comparison coverage, not an independent reproduction of cRia or an official leaderboard submission.

WikiText-2 raw test continuation, reported separately: 291 nonoverlapping blocks of 1024 tokens, 512-token prefix plus 512-token scored suffix; 148,992 scored tokens, 335 tail tokens excluded. GPU evaluation; percentile block bootstrap with 10,000 resamples and seed 2026. PPL intervals exponentiate NLL endpoints. This is not rolling or word perplexity.

  • experimental_bfloat16_pll: NLL 8.238186, 95% CI [8.214775, 8.261612]; pseudo-PPL 3782.672, 95% CI [3695.144, 3872.330].

Structured results and provenance, raw task outputs and per-block continuation scores. HF metadata contains author-reported measurements; no verified benchmark badge is claimed. Historical and new measurements refer to identical weight hashes; individual run manifests retain their original dates and source hashes.

Usage

Install requirements.txt; tested with PyTorch 2.11.0 and Transformers 5.17.0. Loading custom code requires trust_remote_code=True. Example on CPU:

from transformers import AutoTokenizer, AutoModelForMaskedLM
repo = "d0rj/diffusion-51M-base-v1-failed"
tokenizer = AutoTokenizer.from_pretrained(repo)
model = AutoModelForMaskedLM.from_pretrained(repo, trust_remote_code=True).eval()
# Diagnostic only: this failed checkpoint does not produce useful text.
ids = model.generate_masked(batch_size=1, seq_len=16, steps=16, temperature=0.)
print(tokenizer.decode(ids[0], skip_special_tokens=True))

After downloading this repository, install the evaluation dependencies and authenticate for the gated BananaMind dataset after accepting its terms:

pip install -r evaluation/repro/requirements.txt
hf auth login
python evaluation/repro/run.py --device cuda:0 --dtype bfloat16 --batch-size 1 --output comparison-rerun
python evaluation/repro/continuation.py --device cuda:0 --dtype bfloat16 --output continuation-bf16.json

--limit / --limit-blocks are smoke checks only. Dataset examples are not redistributed. The bundled runners use this published model and tokenizer.

Training logs and limitations

TensorBoard event files contain available training telemetry and the full evaluation at optimizer step 15,000. train/loss has 750 visible points spanning steps 20–15,000. Resume/purge records are retained. Comet experiment. Evaluation exports retain numeric metrics, uncertainty and task/subtask breakdowns.

Equal source-token budgets do not imply equal target supervision, parameter count or FLOPs. Training dataset file revisions and overlap with benchmarks were not independently audited. Results come from one training seed; uncertainty intervals do not capture training-seed variability or dependence between templated examples. These are small base models, not instruction-tuned assistants.

Downloads last month
-
Safetensors
Model size
51.4M params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Dataset used to train d0rj/diffusion-51M-base-v1-failed

Collection including d0rj/diffusion-51M-base-v1-failed

Evaluation results