Language Diffusion Models are Associative Memories Capable of Retrieving Unseen Data

Model checkpoints for the paper:

Language Diffusion Models are Associative Memories Capable of Retrieving Unseen Data Bao Pham, Mohammed J. Zaki, Luca Ambrogioni, Dmitry Krotov, Matteo Negri Accepted to EMNLP 2026 (Main Conference). arXiv:2604.26841 · paper · code

Abstract

When do language diffusion models memorize their training data, and how to quantitatively assess their true generative regime? We address these questions by showing that Uniform-based Discrete Diffusion Models (UDDMs) fundamentally behave as Associative Memories (AMs) with emergent creative capabilities. The core idea of an AM is to reliably recover stored data points as memories by establishing distinct basins of attraction around them. Historically, models like Hopfield networks use an explicit energy function to guarantee these stable attractors. We broaden this perspective by leveraging the observation that energy is not strictly necessary, as basins of attraction can also be formed via conditional likelihood maximization. By evaluating token recovery of training and test examples, we identify in UDDMs a sharp memorization-to-generalization transition governed by the size of the training dataset: as it increases, basins around training examples shrink and basins around unseen test examples expand, until both later converge to the same level. Crucially, we can detect this transition using only the conditional entropy of predicted token sequences: memorization is characterized by vanishing conditional entropy, while in the generalization regime the conditional entropy of most tokens remains finite. Thus, conditional entropy offers a practical probe for the memorization-to-generalization transition in deployed models.

What is released here

The training-set size is the axis the paper's transition is measured along, so this repo is a full sweep over it: three model sizes × 54 nested subsets of LM1B, from 0.01% of the corpus to 100%, every checkpoint trained for exactly 1,000,000 steps. 162 checkpoints, ~473 GB.

Repository layout

Checkpoints are grouped into one directory per model size:

tiny/     54 checkpoints    ~0.38 GB each
small/    54 checkpoints    ~2.23 GB each
medium/   54 checkpoints    ~6.15 GB each

File names keep their original form:

<size>/lm1b-<size>-<subset>.ckpt

where <subset> is the fraction of the LM1B training set the model saw (data.subset in the training config). The sweep covers 54 values:

Range Values
Ultra-low data 0.0001
Fine grid 0.0007190.009381 (15 steps of ~0.000619)
Low-data grid 0.01, 0.02, 0.03, 0.04, 0.05, 0.06, 0.07
Main grid 0.1, 0.13, 0.16, … 1.0 (steps of 0.03)

So medium/lm1b-medium-0.0001.ckpt is the medium model trained on 0.01% of LM1B, and medium/lm1b-medium-1.0.ckpt is the same architecture on the full corpus. Holding the architecture fixed and sweeping subset is what isolates the memorization behaviour.

Model sizes

Size Backbone hidden_size n_blocks n_heads cond_dim Params (backbone)
tiny ddit 256 8 8 128 23.7 M
small ddit 768 12 12 128 139.3 M
medium ddit 1024 24 16 128 384.0 M

Shared across all sizes: length: 1024, dropout: 0.1, scale_by_sigma: True, tie_word_embeddings: False, vocab_lookup: True, log-linear noise schedule.

Checkpoint contents

These are full PyTorch Lightning checkpoints, not weights-only exports. Each file contains:

Key Share of file Notes
state_dict ~25% Live backbone weights, fp32
ema ~25% {decay, num_updates, shadow_params} — EMA shadow weights, use these for sampling / evaluation
optimizer_states ~50% AdamW moments; keep these to resume training
loops, callbacks, lr_schedulers, hyper_parameters, sampler <1% Lightning bookkeeping

Because optimizer state is preserved, any checkpoint here can be resumed, not just evaluated.

Usage

Download a single checkpoint:

from huggingface_hub import hf_hub_download

path = hf_hub_download(
    repo_id="lemoncmd/lldms-associative-memory",
    filename="tiny/lm1b-tiny-0.01.ckpt",
)

Download one whole size:

from huggingface_hub import snapshot_download

snapshot_download(
    repo_id="lemoncmd/lldms-associative-memory",
    allow_patterns="medium/*",
)

Load the EMA weights for evaluation:

import torch

ckpt = torch.load(path, map_location="cpu", weights_only=False)
ema = ckpt["ema"]                  # {"decay", "num_updates", "shadow_params"}
shadow = ema["shadow_params"]      # list of tensors, ordered as model.parameters()
train_weights = ckpt["state_dict"] # live (non-EMA) weights
step = ckpt["global_step"]

Note that loading requires TORCH_FORCE_NO_WEIGHTS_ONLY_LOAD=1 on recent PyTorch, or weights_only=False as above, since the checkpoints contain pickled config objects.

Training setup

Trained with the config in the accompanying codebase (Hydra), tokenizer bert-base-uncased (vocab 30,522):

  • Global batch size 512, sequence length 1024
  • Log-linear noise schedule, duo_base algorithm
  • Constant LR with warmup
  • All checkpoints released here are at global_step = 1,000,000
  • DDP across 4× H100 80GB per run

Reproduce a single run with:

python main.py model=medium data.subset=0.25

Citation

@misc{pham2026languagediffusionmodelsassociative,
      title={Language Diffusion Models are Associative Memories Capable of Retrieving Unseen Data},
      author={Bao Pham and Mohammed J. Zaki and Luca Ambrogioni and Dmitry Krotov and Matteo Negri},
      year={2026},
      eprint={2604.26841},
      archivePrefix={arXiv},
      primaryClass={cs.LG},
      url={https://arxiv.org/abs/2604.26841},
}
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Dataset used to train lemoncmd/lldms-associative-memory

Paper for lemoncmd/lldms-associative-memory