Twitter/X

Author claim: Almost every agent-memory system publishes scores on its own…

Brief

Agent Memory Leaderboard is an open, neutral championship (submissions close Aug 7; results mid‑August) requiring entrants to expose only add+search endpoints while organizers run identical generation, judging, and compute and publish logs/evidence. It unifies a 10+ dataset text track (~150M chars, ~5,000 questions, seven capability dimensions) and a code track (12 repos, 150 tasks, 1,290 annotated PRs) with strict temporal constraints.

Why it matters

Author claim: Almost every agent-memory system publishes scores on its own datasets, models, and judges, making it impossible to tell whether higher scores reflect better memory or weaker evaluations.

Key details

  • Agent Memory Leaderboard (AgentMemoryL) runs an open, neutral championship where every entrant exposes only two endpoints (add + search); organizers run identical generation, judging, aggregation and compute, publish logs/evidence, submissions close Aug 7, results mid‑August.
  • Competition specifics: Text track consolidates 10+ datasets (PersonaMem, LoCoMo-Refined, CLBench, BEAM, LongMemEval, ScriptMem, etc.) with ~150M characters of history, ~5,000 questions, scored across 7 capability dimensions; Code track uses 12 repos, 150 base tasks and 1,290 historical PRs annotated, with strict temporal constraints; two boards separate open-source (eligible for rewards) and commercial products; organized by 20+ universities/institutions.
Source evidence

Everyone in agent memory publishes numbers on their own datasets, their own models, their own judges. So when a score goes up, you can't tell if the memory improved or the eval got softer.

@AgentMemoryL is running the first open, neutral championship:

> Every entrant hits the same add + search endpoints
> They handle generation, judging, and compute
> Text track across 10+ datasets, plus a real code track with temporal constraints
> Logs and evidence are public

Same setup for everyone. That's the part that's been missing.

Agent Memory Leaderboard (@AgentMemoryL)

Every agent memory system publishes its own numbers. Almost none of them are comparable.

Different datasets, different answer models, different judges. When a score moves, you can’t tell whether the memory got better or the setup got kinder.

The Agent Memory Challenge has been open since Wednesday.
Submissions close August 7.

The design, briefly:

You expose two endpoints — add and search. That’s the entire integration.
Answer generation, judging, aggregation and orchestration run on our side, identically for every entry, and we pay for that compute. Your internals stay closed.

Text track
10+ datasets consolidated (PersonaMem, LoCoMo-Refined, CLBench, BEAM, LongMemEval, ScriptMem and others).
~150M characters of history, ~5,000 questions, scored across seven capability dimensions rather than collapsed into one number.

Code track — the part we think doesn’t exist yet:
12 repos, 150 base tasks, and 1,290 historical PRs annotated as strongly related, weakly related or unrelated to the task at hand.

Strict temporal constraints, so nothing from the future leaks backward.

It measures whether a system can retrieve and reuse engineering experience, not whether it can solve an issue cold.

Two boards
Open-source methods — ranked and eligible for rewards
Commercial products — ranked separately for procurement reference

Organized by 20+ universities and research institutions.

Enter → agentmemories.ai/competition…
Protocol, submission guide, reference integrations →
github.com/AML-memory/agent-…

Results go up mid-August. Logs, retrieved evidence and judge records stay public.

— https://nitter.net/AgentMemoryL/status/2083234552844337229#m