ArXiv

LongMemEval-V2: Evaluating Long-Term Agent Memory Toward Experienced Colleagues

Authors
Di Wu, Zixiang Ji, Asmi Kawatkar...
Categories
cs.CL
arXiv
https://arxiv.org/abs/2605.12493v1
PDF
https://arxiv.org/pdf/2605.12493v1

Brief

LongMemEval-V2 (LME-V2) is a benchmark for assessing whether memory systems let agents acquire environment-specific experience; it provides 451 questions across five memory abilities with histories up to 500 trajectories (115M tokens). The authors evaluate a RAG-style AgentRunbook-R and a file+coding-agent AgentRunbook-C, reporting 72.5% accuracy for AgentRunbook-C versus 48.5% for the best RAG baseline and 69.3% for a coding-agent baseline, while noting higher latency; summary based on the abstract.

Cleaned source text

Abstract

Comment: Work in Progress