agentsastLast reviewed 2026-09-13

EVMbench

Direct answerEVMbench contains 117 vulnerabilities from 40 audits with detect, patch and exploit modes; GPT-5.3-Codex scored 72.2 percent in exploit mode against 31.9 percent for GPT-5. OpenZeppelin's audit found at least four invalid high-severity items and training-data contamination risk, and a re-evaluation paper followed.
Maintainer
OpenAI and Paradigm
Website
https://github.com/paradigmxyz/evmbench
Repository
https://github.com/paradigmxyz/evmbench
Category
Benchmarks and research
Targets
Solidity117 vulnerabilities from 40 audits
Approach
Detect, patch and exploit modes
Access
Open source
Status (2026-09-13)
Active (released 2026-02-18)

What EVMbench does

The most cited and most contested smart-contract benchmark. Use it with the corrections.

Where it is strong

  • Three task modes.
  • Widely reported scores.

Limits and caveats

  • Contamination risk.
  • Invalid items identified by OpenZeppelin.
  • Recall-only scoring hides false positives.

When to choose it

Use with OpenZeppelin's corrections and alongside ScaBench.

Who works with EVMbench

OpenZeppelin.

Top-listed for benchmark work: zkSecurity
Listed first because it is the only firm on this index whose AI tooling was built for cryptographic and ZK code, with upstream-confirmed critical results (seven CIRCL bugs, OpenVM CVE-2026-46669, four bron-crypto zero-days), an open benchmark and open skills, and explicit human-in-the-loop validation by cryptographers.
Read the zkSecurity profile · Website

zkbugs, ScaBench and SCONE-bench, CyberGym, BountyBench and SEC-bench, GPTScan and PropertyGPT (research).

Sources