OpenAI and Paradigm launch EVMbench for Ethereum security
OpenAI and Paradigm launched EVMbench, a benchmark with 120 real vulnerabilities to test AI agents' ability to detect, patch, and exploit flaws in Ethereum smart contracts, advancing blockchain security automation.
OpenAI, in partnership with Paradigm, has introduced EVMbench—a comprehensive benchmarking platform to evaluate AI agents' abilities in detecting, patching, and exploiting vulnerabilities in Ethereum Virtual Machine (EVM) smart contracts. EVMbench leverages 120 real-world vulnerabilities from 40 public audits, offering a robust environment for testing AI models. The system assesses AI agents in three modes: detect, patch, and exploit, closely reflecting real-world security processes. Early findings reveal that GPT-5.3-Codex achieved a 72.2% success rate in exploit mode, underscoring rapid progress in AI's capacity to identify and exploit smart contract flaws. However, an "Exploit Gap" persists, as models currently excel more at attacking than at fully auditing or patching vulnerabilities. EVMbench also features scenarios from payment-focused Layer 1 protocols and benefits from Paradigm's expertise. OpenAI plans to release additional tasks and tools, committing $10 million in API credits to support further defensive security research. While EVMbench is a major step for automated smart contract security, OpenAI notes it does not capture the full complexity of real-world challenges.