Abstract
Can AI agents conduct open-ended AI research? Early evidence from two case studies
- Authors
- Peter Kirgis, Sayash Kapoor, Andrew Schwartz...
- Categories
- cs.AI, cs.CY, cs.LG
Brief
The paper introduces shadow evaluations—having frontier agents tackle the central open-ended question of high-quality unpublished papers and having the original authors grade outputs—and applies it to two NeurIPS 2026 submissions. Despite six days and substantial compute, agents handled engineering but failed to advance the research questions, revealing five systematic failure modes; artifacts and reviews are publicly released.
Why it matters
Agents ran shadow evaluations on two unpublished NeurIPS 2026 submissions, given six days and thousands of dollars of compute; agents completed all engineering tasks without human help but made no substantial progress on the core research questions, and both papers were unambiguously rejected by their original authors.
Key details
- The study identifies five recurring failure modes—poor judgment about the bar for publishable research, uncreative fixes to research-design shortcomings, ineffective backtracking from dead ends, poor resource awareness, and instruction drift—and a robustness check with a second model and scaffold reproduced these failures.
- The authors introduce the 'shadow evaluation' method for measuring AI R&D automation and release expert reviews, survey responses, agent repositories, and logs to support reproducibility and further analysis.