agentsastLast reviewed 2026-09-13

Benchmark contamination

Direct answerThe model has seen the audit report or the bug in training, so a benchmark hit measures recall of memory rather than discovery.

In more detail

OpenZeppelin raised this about EVMbench; it applies to every benchmark built from public audits, including zkbugs. The defences are benchmarks that post-date model cutoffs, private held-out sets, and full-codebase modes that test search rather than recognition.

Tools that address it

EVMbench, zkbugs.

False positive rate, Precision vs recall, Agentic scanning, LLM plus fuzzing, LLM plus symbolic execution or formal verification, Hallucinated vulnerabilities, Triage burden, Human-in-the-loop, AI-assisted audit vs AI audit, Prompt injection in auditing pipelines, Responsible disclosure of AI-found bugs, Continuous scanning and run-count coverage, Proof-of-concept harness, Threat model file, Severity calibration

Getting help

Firms on this index that handle this in practice: zkSecurity, Trail of Bits, Zellic, Nethermind Security.