Your AGENTS.md may be silently wasting agent steps. Probe-and-refine tunes repo guidance with synthetic bug-fix probes: 33.0% SWE-bench Verified vs 28.3% for a static KB. The gain is coverage: more evaluable patches, same precision. arxiv.org/abs/2606.20512
Link
Probe-and-Refine Tuning of Repository Guidance for Coding Agents
LLM-based coding agents need higher-level operational knowledge about a repository (which files house which subsystems, how to run the test suite, which workflows have historically led to wrong...
arxiv.org