Twitter/X

Boundary-Bench (paper, benchmark, and GitHub repo) was open-sourced on 2026-08-05…

Brief

Boundary-Bench is an open-sourced paper, benchmark, and GitHub repo (released 2026-08-05) from Accomplish_ai and NYU that ran ~10,000 tests on 12 frontier agents in simulated enterprise settings with EDR, SASE, and DLP enforcement. The project aims to expose how real-world security policies change agent performance versus standard leaderboard results.

Why it matters

Boundary-Bench (paper, benchmark, and GitHub repo) was open-sourced on 2026-08-05 by @_orcaman; the project was developed by researchers at Accomplish_ai and NYU.

Key details

  • The benchmark evaluated 12 frontier agents across roughly 10,000 runs in simulated enterprise environments with EDR, SASE, and DLP tools enforcing realistic policies.
  • Boundary-Bench claims agent performance under realistic security controls diverged substantially from generic leaderboard scores, implying organizations may be making deployment and risk decisions based on non-representative metrics.
Source evidence

Today we are open-sourcing @boundarybench, a new paper, benchmark, and GitHub repo that enterprises can use to find out the REAL performance of their agents.

Boundary-Bench was developed by researchers from @Accomplish_ai and NYU, where we tested 12 frontier agents across roughly 10,000 runs, with realistic enterprise policies, simulating environments with EDR, SASE, and DLP security tools enforcing those policies.

We did this because generic leaderboard scores are being generated under conditions no security team would ever allow, which means orgs are making deployment and risk decisions based on numbers that don't hold up.

The results are surprising >>

Video