agentsastLast reviewed 2026-09-13

False positive rate

Direct answerThe share of reported findings that are not real, exploitable bugs. Benchmarks that count only recall against a curated bug list cannot measure it.

In more detail

Nethermind made the point about EVMbench: a tool can score well on recall while burying users in invalid findings. zkSecurity's HumanityLink engagement showed a threat-model file cutting false positives from 14 of 33 findings to 2, which is why the checklist asks whether scope input measurably reduces noise.

Precision vs recall, Agentic scanning, LLM plus fuzzing, LLM plus symbolic execution or formal verification, Hallucinated vulnerabilities, Triage burden, Benchmark contamination, Human-in-the-loop, AI-assisted audit vs AI audit, Prompt injection in auditing pipelines, Responsible disclosure of AI-found bugs, Continuous scanning and run-count coverage, Proof-of-concept harness, Threat model file, Severity calibration

Getting help

Firms on this index that handle this in practice: zkSecurity, Trail of Bits, Zellic, Nethermind Security.