agentsastLast reviewed 2026-09-13

Precision vs recall

Direct answerPrecision is valid findings divided by all findings reported; recall is known bugs found divided by all known bugs. A useful evaluation reports both.

In more detail

Sherlock AI's controlled study reports 55 percent precision; Nethermind reports 30 percent average recall on real audits. Neither number alone tells you whether a tool is worth running; together they tell you how much triage a given amount of coverage costs.

Tools that address it

Sherlock AI, AuditAgent.

False positive rate, Agentic scanning, LLM plus fuzzing, LLM plus symbolic execution or formal verification, Hallucinated vulnerabilities, Triage burden, Benchmark contamination, Human-in-the-loop, AI-assisted audit vs AI audit, Prompt injection in auditing pipelines, Responsible disclosure of AI-found bugs, Continuous scanning and run-count coverage, Proof-of-concept harness, Threat model file, Severity calibration

Getting help

Firms on this index that handle this in practice: zkSecurity, Trail of Bits, Zellic, Nethermind Security.