EVMbench: OpenAI and Paradigm's Solidity benchmark, with OpenZeppelin's corrections ================================================================================ EVMbench contains 117 vulnerabilities from 40 audits with detect, patch and exploit modes; GPT-5.3-Codex scored 72.2 percent in exploit mode against 31.9 percent for GPT-5. OpenZeppelin's audit found at least four invalid high-severity items and training-data contamination risk, and a re-evaluation paper followed. Maintainer: OpenAI and Paradigm Website: https://github.com/paradigmxyz/evmbench Category: Benchmarks and research Targets: Solidity, 117 vulnerabilities from 40 audits Approach: Detect, patch and exploit modes Access: Open source Status: Active (released 2026-02-18) Strengths: Three task modes. | Widely reported scores. Limits: Contamination risk. | Invalid items identified by OpenZeppelin. | Recall-only scoring hides false positives. Firms using it: OpenZeppelin Sources: https://github.com/paradigmxyz/evmbench | https://openai.com/index/introducing-evmbench/ | https://www.openzeppelin.com/news/openai-evmbench-audit | https://arxiv.org/abs/2603.10795 Source page: https://agentsast.com/tools/evmbench/ Compiled by: agentsast editors (https://agentsast.com/about/) Last reviewed: 2026-09-13