Today we are open-sourcing @boundarybench, a new paper, benchmark, and GitHub repo that enterprises can use to find out the REAL performance of their agents.
Boundary-Bench was developed by researchers from @Accomplish_ai and NYU, where we tested 12 frontier agents across roughly 10,000 runs, with realistic enterprise policies, simulating environments with EDR, SASE, and DLP security tools enforcing those policies.
We did this because generic leaderboard scores are being generated under conditions no security team would ever allow, which means orgs are making deployment and risk decisions based on numbers that don't hold up.
The results are surprising >>
Video