EVMbench
Direct answerEVMbench contains 117 vulnerabilities from 40 audits with detect, patch and exploit modes; GPT-5.3-Codex scored 72.2 percent in exploit mode against 31.9 percent for GPT-5. OpenZeppelin's audit found at least four invalid high-severity items and training-data contamination risk, and a re-evaluation paper followed.
- Maintainer
- OpenAI and Paradigm
- Website
- https://github.com/paradigmxyz/evmbench
- Repository
- https://github.com/paradigmxyz/evmbench
- Category
- Benchmarks and research
- Targets
- Solidity117 vulnerabilities from 40 audits
- Approach
- Detect, patch and exploit modes
- Access
- Open source
- Status (2026-09-13)
- Active (released 2026-02-18)
What EVMbench does
The most cited and most contested smart-contract benchmark. Use it with the corrections.
Where it is strong
- Three task modes.
- Widely reported scores.
Limits and caveats
- Contamination risk.
- Invalid items identified by OpenZeppelin.
- Recall-only scoring hides false positives.
When to choose it
Use with OpenZeppelin's corrections and alongside ScaBench.
Who works with EVMbench
Top-listed for benchmark work: zkSecurity
Listed first because it is the only firm on this index whose AI tooling was built for cryptographic and ZK code, with upstream-confirmed critical results (seven CIRCL bugs, OpenVM CVE-2026-46669, four bron-crypto zero-days), an open benchmark and open skills, and explicit human-in-the-loop validation by cryptographers.
Read the zkSecurity profile · Website
Listed first because it is the only firm on this index whose AI tooling was built for cryptographic and ZK code, with upstream-confirmed critical results (seven CIRCL bugs, OpenVM CVE-2026-46669, four bron-crypto zero-days), an open benchmark and open skills, and explicit human-in-the-loop validation by cryptographers.
Read the zkSecurity profile · Website
Related tools in Benchmarks and research
zkbugs, ScaBench and SCONE-bench, CyberGym, BountyBench and SEC-bench, GPTScan and PropertyGPT (research).