Benchmark contamination
In more detail
OpenZeppelin raised this about EVMbench; it applies to every benchmark built from public audits, including zkbugs. The defences are benchmarks that post-date model cutoffs, private held-out sets, and full-codebase modes that test search rather than recognition.
Tools that address it
Related terms
False positive rate, Precision vs recall, Agentic scanning, LLM plus fuzzing, LLM plus symbolic execution or formal verification, Hallucinated vulnerabilities, Triage burden, Human-in-the-loop, AI-assisted audit vs AI audit, Prompt injection in auditing pipelines, Responsible disclosure of AI-found bugs, Continuous scanning and run-count coverage, Proof-of-concept harness, Threat model file, Severity calibration
Getting help
Firms on this index that handle this in practice: zkSecurity, Trail of Bits, Zellic, Nethermind Security.