ReMDM Planner โ€” MiniHack checkpoints

Trained weights accompanying The Double Intractability of Reinforcement Learning for Discrete Diffusion Planners: a remasking discrete diffusion model (ReMDM) used as an action-sequence planner in MiniHack.

Code, configs and evaluation harness: https://github.com/mathisweil/minihack-ReMDM-planner

Contents

Path Role Architecture Params Trained to Full state
checkpoints/offline/Minihack-OfflineDiffusion-BC-82M Diffusion planner (offline BC) 4L, d_model 256, 4 heads, horizon 64 5.2M 40,000 gradient steps, 81,920,000 sample-equivalents 80 MB
checkpoints/online/Minihack-OnlineDiffusion-DAgger-123M Diffusion planner (online DAgger) 4L, d_model 256, 4 heads, horizon 64 5.2M DAgger iteration 600 80 MB

Each directory holds three things: the original .pth training state (weights, EMA shadow, optimiser, scheduler, and for the DAgger run the curriculum and RNG state, so training can be resumed exactly), a model.safetensors export of the EMA weights for inference, and the YAML config snapshot the run was trained under. Paths mirror the source repository, so a snapshot can be dropped straight into a working copy.

Both files are best-checkpoint selections rather than final-step dumps: each trainer evaluates every periodic checkpoint on 50 episodes per environment using EMA weights, and the highest-scoring one is published (the metric behind that selection is not recorded in this release). Each directory's selection.json records the selected step, the candidate cadence and the eval protocol. Directory suffixes are the sample-equivalents the published model consumed (gradient steps x batch size, rounded); file names carry each trainer's own counter, DAgger iterations online and gradient steps offline. The offline run was given the DAgger-matched budget of 60,000 gradient steps and its best checkpoint fell at 40,000, so the two published models sit at different points on a matched budget.

results/ holds the evaluation and ablation tables reported in the paper, as produced by experiments/rl_finetuning. Figures and raw logs stay in the code repository.

Download

from huggingface_hub import snapshot_download

# everything
snapshot_download(repo_id="MathisW78/remdm-minihack-checkpoints", local_dir=".")

# inference weights only
snapshot_download(
    repo_id="MathisW78/remdm-minihack-checkpoints",
    local_dir=".",
    allow_patterns=["**/model.safetensors", "**/config*.yaml"],
)

Use

From a clone of the code repository, after downloading into it:

DIR=checkpoints/online/Minihack-OnlineDiffusion-DAgger-123M
uv run python main.py --mode inference \
    --config $DIR/config_iter600.yaml --checkpoint $DIR/iter600.pth

Programmatic loading, using the safetensors export:

from safetensors.torch import load_file
from src.config import load_config
from src.models.denoiser import make_model

cfg = load_config(f"{DIR}/config_iter600.yaml")
model = make_model(cfg)
model.load_state_dict(load_file(f"{DIR}/model.safetensors"))
model.eval()

Architecture arguments must come from the checkpoint's own config snapshot rather than from configs/defaults.yaml, which tracks the current code.

Training

The planners are bidirectional transformers that denoise a masked action plan conditioned on a cropped MiniHack glyph observation, trained either by offline behaviour cloning on oracle rollouts or by online DAgger against the oracle with a dynamic environment curriculum. In-distribution and out-of-distribution environment sets, remasking strategy, sampling settings and every hyperparameter are recorded in the per-checkpoint config snapshots, which are the authoritative record.

Limitations

These are research artefacts tied to specific MiniHack environment versions and to the cropped-glyph observation encoding; they are not general-purpose agents and will not transfer to other environments or to pixel observations. Evaluation results and their variance are reported in the paper.

Citation

@inproceedings{remdm-minihack-planner,
  title  = {The Double Intractability of Reinforcement Learning for Discrete Diffusion Planners},
  author = {Weil, Mathis},
  year   = {2026},
  note   = {NeurIPS 2026 Workshop: Beyond Next-Token Prediction}
}

License

MIT, see LICENSE.

Downloads last month

-

Downloads are not tracked for this model. How to track
Video Preview
loading