Twitter/X

PAST-Bench provides 26 task-family scenarios and 204 episodes, evaluates four…

Brief

PAST-Bench introduces a 26‑scenario, 204‑episode benchmark measuring whether retained experience improves personal agents across Memory, Procedural reuse, Information gathering, and Update. Evaluated on 7 base models and 4 frameworks with matched persistence-on/off runs, it finds performance lifts of +0.13–+0.24 and provides mechanism-level attribution. The Hermes+ architecture (five diagnosis-driven interventions) increases the persistence gap and mechanism-evidence (0.64→0.73), especially on Update.

Why it matters

PAST-Bench provides 26 task-family scenarios and 204 episodes, evaluates four capabilities (Memory, Procedural reuse, Information gathering, Update), uses matched persistence-on/off evaluation, and tests 7 base models across 4 agent frameworks.

Key details

  • Retained experience improved downstream task performance by between +0.13 and +0.24 depending on the model, but gains are capability-, model-, and framework-dependent; the benchmark attributes which persistence mechanisms produced those gains.
  • Hermes+ is an architecture with five diagnosis-driven interventions that raises the persistence gap from +0.13 to +0.15, increases the mechanism-evidence score from 0.64 to 0.73, and shows its clearest improvement on the Update capability.
Source evidence

Great read on experience-driven AI agents. The Hermes+ architecture is worth a look if you're building persistent agents.

The paper introduces PAST-Bench, a benchmark for measuring whether retained experience actually improves an agent's future behavior.

It evaluates four capabilities:
• Memory
• Procedural reuse
• Information gathering
• Updates

Across 7 frontier models, retained experience improved performance by +0.13 to +0.24 depending on the model.

What I found interesting is that it doesn't stop at benchmark scores. It also attributes where those gains came from, making it possible to distinguish agents with similar performance but different persistence mechanisms.

Ling Yang (@LingYang_PU)

1/5 🚀 Introducing PAST-Bench: Benchmarking Experience-Driven Evolution in Personal Agents

Do personal agents actually improve from past experience—or do they simply accumulate more state?

PAST-Bench provides:

📚 26 task-family scenarios and 204 episodes
🧠 4 capabilities: Memory, Procedural Reuse, Information Gathering, and Update
⚖️ Matched persistence-on/off evaluation
🔍 Task performance + mechanism-level attribution
🤖 Evaluation across 7 base models and 4 agent frameworks
🛠️ Hermes+: 5 diagnosis-driven interventions across the agent loop

Our key finding: retained experience can improve future behavior, but the gains are highly capability-, model-, and framework-dependent. Agents with the same headline improvement may rely on very different persistence pathways.

Hermes+ improves the persistence gap from +0.13 to +0.15 and the mechanism-evidence score from 0.64 to 0.73, with its clearest gain on Update.

📄 Paper: arxiv.org/abs/2608.04003
💻 Code: github.com/Gen-Verse/PAST-Be…

AIAgents #LLM #AgentMemory #SelfEvolvingAgents #AIResearch #RSI #RecursiveAI

— https://nitter.net/LingYang_PU/status/2084817609674141939#m