Decentralized Master-Mind: Joint Action Refinement through Iterative Intent Denoising in Multi-Agent Pathfinding
Paper • 2609.32019 • Published • 41
DMM is a decentralized multi-agent pathfinding policy that refines agents' action intents over several local communication rounds before committing to actions. The models are pretrained on expert solutions with imitation learning and optionally fine-tuned with MICPO, a critic-free reinforcement learning method.
| Checkpoint | Imitation pretraining iterations | MICPO optimizer updates |
|---|---|---|
| DMM-08M.pt | 1,000,000 | — |
| DMM-3M.pt | 1,000,000 | — |
| DMM-MICPO-08M.pt | 1,000,000 | 96,000 |
| DMM-MICPO-3M.pt | 1,000,000 | 96,000 |
MICPO fine-tuning comprises 500 outer iterations (96,000 optimizer updates). Both model sizes use four communication rounds. Training details are reported in the paper.
See the GitHub repository for code, usage instructions, training, and evaluation.