Instructions to use d0rj/diffusion-51M-base-v1-failed with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use d0rj/diffusion-51M-base-v1-failed with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("fill-mask", model="d0rj/diffusion-51M-base-v1-failed", trust_remote_code=True)# Load model directly from transformers import AutoModelForMaskedLM model = AutoModelForMaskedLM.from_pretrained("d0rj/diffusion-51M-base-v1-failed", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
Diffusion v1 (failed; PLL)
Failed ablation: token-frequency/context-collapse behavior. Retained for reproducibility and negative results. The repaired model is diffusion v2.
A 51,392,512-parameter English base model from the Tiny llm ablation experiment. Trained from random initialization for 15,000 optimizer steps, processing 3,932,160,000 source tokens. These are processed input tokens, not unique text or supervised target-token counts.
Architecture
Failed v1 absorbing-mask diffusion ablation: 10 bidirectional layers, width 512, SwiGLU width 1536, 8 attention heads, head dimension 64, RoPE, RMSNorm and tied 32,768-token output vocabulary. A separate learned mask input vector uses ID 32768, outside the output vocabulary. Legacy additive time conditioning is retained exactly. The checkpoint stores 51,392,512 parameters.
Tokenizer foundation: q-project/Q-50M-Base. This model is trained from scratch, not fine-tuned from the reference weights. Exact trained configuration and custom model code are included.
Training
- Data: FineWeb-Edu,
sample-10BT, local Parquet shards; shuffle buffer 100,000. - Objective: Sample noise time uniformly; independently mask tokens with that probability, predict masked originals with inverse-time weighting normalized by source token count. This run collapsed toward token-frequency predictions and is published as a failed ablation. The repaired diffusion v2 disables additive time conditioning.
- Final recorded batch: 8 sequences × 16 gradient accumulation × 2048 tokens = 262,144 source tokens per update.
- Fused AdamW, peak LR 0.001, betas (0.9, 0.95), weight decay 0.1 excluding bias/norm/1D parameters, gradient clipping 1.0. Warmup 150 steps, cosine decay to 10% of peak LR.
- BF16 compute, one RTX 5070 Ti, seed 2026; released checkpoint weights remain FP32. Training configuration.
Evaluation
Full selected splits, lm-eval 0.4.12, no added few-shot examples or chat template, BF16 on RTX 5070 Ti, context cap 2048 (ArithMark 1024). New evaluations explicitly disable TF32; historical core manifests predate the explicit TF32 flag. Scores below are percentages; ± is one standard error, and the separate interval column is 95% CI where available.
| Dataset | Split | Examples | Metric | Score ± SE (%) | 95% CI (%) |
|---|---|---|---|---|---|
| HellaSwag | validation | 10,042 | pll_acc_norm |
25.97 ± 0.44 | [25.12, 26.84] |
| ARC-Easy | test | 2,376 | pll_acc_norm |
26.68 ± 0.91 | [24.94, 28.50] |
| ARC-Challenge | test | 1,172 | pll_acc_norm |
24.06 ± 1.25 | [21.70, 26.59] |
| PIQA | validation | 1,838 | pll_acc_norm |
50.05 ± 1.17 | [47.77, 52.34] |
| WinoGrande | validation | 1,267 | pll_acc |
50.99 ± 1.40 | [48.24, 53.73] |
| OpenBookQA | test | 500 | pll_acc_norm |
25.80 ± 1.96 | [22.16, 29.81] |
| BoolQ | validation | 3,270 | pll_acc |
37.83 ± 0.85 | [36.18, 39.50] |
| LAMBADA OpenAI | test | 5,153 | pll_acc |
0.00 ± 0.00 | [0.00, 0.07] |
| ArithMark-3 | train | 1,000 | pll_acc_norm |
29.90 ± 1.45 | [27.14, 32.81] |
| Balanced COPA | train | 1,000 | pll_acc |
50.00 ± 1.58 | [46.91, 53.09] |
| BLiMP | train | 67,000 | pll_acc |
57.39 ± 0.16 | — |
| CommonsenseQA | validation | 1,221 | pll_acc |
19.57 ± 1.14 | [17.45, 21.89] |
| MMLU continuation | test | 14,042 | pll_acc |
23.38 ± 0.36 | — |
| SciQ (with support) | test | 1,000 | pll_acc_norm |
23.70 ± 1.35 | [21.17, 26.43] |
| TruthfulQA MC2 | validation | 817 | pll_acc |
48.99 ± 1.64 | — |
| BananaMind Base 1.1 | test | 350 | pll_raw_accuracy |
26.29 ± 2.36 | [21.95, 31.14] |
These are experimental single-mask continuation PLL scores, not autoregressive likelihood. Other answer tokens remain visible. LAMBADA is whole-answer reconstruction under this protocol, not ordinary next-word generation. Compare this model in a separate diagnostic group. LAMBADA requires all final-word tokens to match. The descriptive macro-average is 32.5378%: one primary metric per each of the 16 tasks, with equal task weight. PPL, standard errors, CI bounds and MMLU/BLiMP leaf scores are excluded. It is not an official leaderboard score or a statistical ranking test.
Full selected splits; lm-eval 0.4.12; seed 1234; BF16 on RTX 5070 Ti; context cap 2048 (ArithMark 1024), TF32 disabled on new runs (historical core flag unrecorded), no chat template and no added few-shot examples. TruthfulQA retains the harness's fixed six-QA preamble. ArithMark and BananaMind normalize by continuation token count; ordinary harness acc_norm uses its own length normalization. BananaMind is raw accuracy, not Elo. SciQ includes the support passage. Balanced COPA uses the mirrored 1000-item train-named evaluation split; cRia's split was inferred, not confirmed. MMLU scores full answer continuations across 57 subjects, weighted by item count; BLiMP averages 67 equal-sized minimal-pair subsets. Standard errors are retained from each evaluator. Wilson intervals are reported only where the runner logged binary item accuracy; MC2 is probability mass, not binary accuracy. These intervals do not model dependence between paired/templated examples or training seed variation. UL2, PrefixLM and experimental diffusion PLL use their documented conditional scoring protocols; PLL exposes the other answer tokens and is not autoregressive likelihood. cRia's published scores used a different precision and benchmark-adapted checkpoint; this completes our comparison coverage, not an independent reproduction of cRia or an official leaderboard submission.
WikiText-2 raw test continuation, reported separately: 291 nonoverlapping blocks of 1024 tokens, 512-token prefix plus 512-token scored suffix; 148,992 scored tokens, 335 tail tokens excluded. GPU evaluation; percentile block bootstrap with 10,000 resamples and seed 2026. PPL intervals exponentiate NLL endpoints. This is not rolling or word perplexity.
experimental_bfloat16_pll: NLL 8.238186, 95% CI [8.214775, 8.261612]; pseudo-PPL 3782.672, 95% CI [3695.144, 3872.330].
Structured results and provenance, raw task outputs and per-block continuation scores. HF metadata contains author-reported measurements; no verified benchmark badge is claimed. Historical and new measurements refer to identical weight hashes; individual run manifests retain their original dates and source hashes.
Usage
Install requirements.txt; tested with PyTorch 2.11.0 and Transformers 5.17.0. Loading custom code requires trust_remote_code=True. Example on CPU:
from transformers import AutoTokenizer, AutoModelForMaskedLM
repo = "d0rj/diffusion-51M-base-v1-failed"
tokenizer = AutoTokenizer.from_pretrained(repo)
model = AutoModelForMaskedLM.from_pretrained(repo, trust_remote_code=True).eval()
# Diagnostic only: this failed checkpoint does not produce useful text.
ids = model.generate_masked(batch_size=1, seq_len=16, steps=16, temperature=0.)
print(tokenizer.decode(ids[0], skip_special_tokens=True))
After downloading this repository, install the evaluation dependencies and authenticate for the gated BananaMind dataset after accepting its terms:
pip install -r evaluation/repro/requirements.txt
hf auth login
python evaluation/repro/run.py --device cuda:0 --dtype bfloat16 --batch-size 1 --output comparison-rerun
python evaluation/repro/continuation.py --device cuda:0 --dtype bfloat16 --output continuation-bf16.json
--limit / --limit-blocks are smoke checks only. Dataset examples are not redistributed. The bundled runners use this published model and tokenizer.
Training logs and limitations
TensorBoard event files contain available training telemetry and the full evaluation at optimizer step 15,000. train/loss has 750 visible points spanning steps 20–15,000. Resume/purge records are retained. Comet experiment. Evaluation exports retain numeric metrics, uncertainty and task/subtask breakdowns.
Equal source-token budgets do not imply equal target supervision, parameter count or FLOPs. Training dataset file revisions and overlap with benchmarks were not independently audited. Results come from one training seed; uncertainty intervals do not capture training-seed variability or dependence between templated examples. These are small base models, not instruction-tuned assistants.
- Downloads last month
- -
Dataset used to train d0rj/diffusion-51M-base-v1-failed
Collection including d0rj/diffusion-51M-base-v1-failed
Evaluation results
- pll_acc_norm (fraction; full comparison protocol) on HellaSwagvalidation set self-reported0.260
- pll_acc_norm (fraction; full comparison protocol) on ARC-Easytest set self-reported0.267
- pll_acc_norm (fraction; full comparison protocol) on ARC-Challengetest set self-reported0.241
- pll_acc_norm (fraction; full comparison protocol) on PIQAvalidation set self-reported0.501
- pll_acc (fraction; full comparison protocol) on WinoGrandevalidation set self-reported0.510
- pll_acc_norm (fraction; full comparison protocol) on OpenBookQAtest set self-reported0.258
- pll_acc (fraction; full comparison protocol) on BoolQvalidation set self-reported0.378
- pll_acc (fraction; full comparison protocol) on LAMBADA OpenAItest set self-reported0.000