TWITTER_POST

OpenAI introduced EVMbench as a benchmark for testing AI agents on EVM smart…

Brief

OpenAI introduced EVMbench as a benchmark for testing AI agents on EVM smart contract security. The benchmark focuses specifically on high-severity vulnerabilities and evaluates a full workflow: identifying flaws, successfully exploiting them, and then producing patches, positioning smart contract auditing as an end-to-end agent capability.

Source evidence

title: @OpenAI: Introducing EVMbench—a new benchmark that measures how well AI agents can detect...
author: OpenAI
contenttype: twitterpost
published: 2026-02-18T18:46:44+00:00
source_url: https://x.com/OpenAI/status/2024193883748651102

word_count: 19

Tweet by @OpenAI

Introducing EVMbench—a new benchmark that measures how well AI agents can detect, exploit, and patch high-severity smart contract vulnerabilities.


Posted: 2026-02-18T18:46:44.000Z
Engagement: 8848 likes, 2125 retweets, 1166 replies