Agent Memory Leaderboard
Unified memory evaluation · Results expected August 12.
Agent memory, long-term memory, memory-augmented agents, retrieval and RAG, evaluation and benchmarking, continual learning, long-context systems, multimodal memory, and LLM agents.
A unified evaluation platform for long-term memory systems and memory-enabled agents.
Agent Memory Leaderboard (a.k.a. 记忆之巅) evaluates how effectively memory systems store, retrieve, and use information across long conversations, persistent user contexts, and memory-intensive agent tasks.
Official evaluation and submission are conducted exclusively through the Agent Memory Leaderboard website. This Hugging Face organization hosts public leaderboard releases, evaluation documentation, and community updates.
| Track | Intended for |
|---|---|
| Academic Methods | Reproducible research systems, open-source methods, and academic implementations |
| Industry Systems | Production APIs, hosted services, commercial systems, and closed-source products |
All systems are evaluated under a versioned evaluation contract with fixed datasets, prompts, answer models, judge configurations, and pipeline hashes.
The inaugural public leaderboard is scheduled for mid-August 2026.
Results published on Hugging Face are official release snapshots. The live and authoritative leaderboard remains on the Agent Memory Leaderboard website.
This organization will publish:
Follow this organization for the inaugural leaderboard release and future evaluation cycles.