DiffuThink-Story-512

A 13,439,680-parameter hybrid language model whose entire weight lineage was trained from random initialization in this project. No pretrained model weights or external tokenizer vocabulary were used. Custom PyTorch architecture and inference API; not an AutoModel-compatible checkpoint.

What it does

Completes English story fragments and reconstructs masked subwords. Uses bidirectional RoPE attention, RMSNorm, SwiGLU, learned absolute positions, tied input/output embeddings and noise conditioning. The training objective mixes random masks and masked suffixes. This hybrid checkpoint adds 90% causal and 10% denoising training batches. Continuation uses strictly causal next-token prediction. Infilling uses confidence-ordered unmasking; autoregressive continuation is NOT diffusion. The optional adaptive budget is a confidence heuristic, not learned reasoning or a new diffusion theorem.

Measured held-out reconstruction

Mask rate Top-1 token accuracy Top-5 Cross-entropy (nats) Unigram top-1
15% 69.39% 87.79% 1.354 7.04%
50% 51.36% 73.87% 2.348 6.89%
85% 23.93% 44.64% 4.233 6.96%

Exact experiment details, counts, versions and limitations: evaluation.json. These numbers are subword reconstruction metrics, not chat quality or autoregressive perplexity. The test uses the first 512 prepared held-out windows.

Run locally

Download this repository, install requirements.txt, and run from its directory:

from diffuthink.v2.inference import load
from diffuthink.v2.hybrid import generate
model, tokenizer = load(".", device="cpu")
result = generate(model, tokenizer,
    "Once upon a time, a little girl found", max_new_tokens=192,
    temperature=0.5, seed=42, precision="fp32", finish_sentence_tokens=32)
print(result["text"])

For a Hub download use huggingface_hub.snapshot_download(repo_id=YOUR_REPO_ID), then use that returned directory as the model path and Python import root. The Python code is included for inspection; no remote-code execution is required by a loader.

Data and provenance

TinyStories: https://huggingface.co/datasets/roneneldan/TinyStories Pinned revision: f54c09fd23315a6f9c86f9dc80f725de7d8f9c64. 500,000 training stories; 1,000 validation and 1,000 test stories from the source validation file. Normalized exact-document duplicates are removed across splits; near-duplicate and semantic leakage detection is not claimed. Each document is windowed only after splitting. The byte-level BPE tokenizer is trained only on training stories (first 50,000). TinyStories is synthetic, generated using GPT-3.5/4, and its card declares CDLA-Sharing-1.0. The data is not included in this model release. See training_info.json for lineage and hashes.

Limitations and authorship

Small English story model: can repeat, invent, stop early, or produce ungrammatical text. No general reasoning, instruction following, factual accuracy, multilingual proficiency, or superiority to established language models is claimed. The basic unigram comparison is weak; sampler and bigram experiments are documented separately in the project. Development was substantially AI-assisted. The portfolio contribution is the implemented pipeline, experimental method, analysis and subsequent personal work, not a claim to have invented Transformers, BPE or discrete diffusion.

Publication

This is a local release candidate, not an uploaded model. Select an appropriate project/model license before a public release and preserve source-data attribution. No license tag is asserted automatically for the trained weights.

Causal evaluation

Held-out next-token NLL: 1.6798 nats; perplexity: 5.364. Evaluated over 95831 targets. This is a genuine causal likelihood metric, unlike masked reconstruction CE. It does not measure story coherence.

Default decoding uses top-p 0.9, repetition penalty 1.12 and no repeated four-token sequences. These are disclosed logit constraints, not grammatical rewriting. The project report also preserves samples without repetition controls.

Expanded context phase

Context: 512 tokens. The original positional embeddings were preserved; added positions were initialized near their mean and then trained. The BPE vocabulary and held-out document splits are unchanged. This is continued training of this project's own from-scratch weights.

The local demo allows up to 32 extra tokens after its requested budget to reach terminal punctuation. It reports whether EOS, punctuation, or the hard length limit ended generation. That display heuristic does not guarantee a finished or coherent story. Fair before/after benchmarks disable the extra-token allowance.

Controlled previous-checkpoint comparison

Both models were evaluated on the same 1380 legacy test windows (192 tokens), with identical targets and tokenizer. Perplexity: 7.206 before, 6.147 after. This differs from the 512-token-window evaluation above. All 24-prompt, two-seed samples and their decoding settings are in comparison.json. Repetition and EOS rates are descriptive indicators, not a semantic-coherence assessment. Corpus size, context and training compute changed together; their individual effects are not isolated.

Downloads last month
142
Safetensors
Model size
13.4M params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Dataset used to train OthmaneW/DiffuThink-Story-512