agentsastLast reviewed 2026-09-13

Severity calibration

Direct answerWhether the tool's assigned severities match what an expert would assign. In the CIRCL study, four of seven AI severities were too high and one critical was rated medium.

In more detail

Mis-calibration in both directions is expected; a human must re-rate. Reports that pass AI severities through unchanged are a red flag.

False positive rate, Precision vs recall, Agentic scanning, LLM plus fuzzing, LLM plus symbolic execution or formal verification, Hallucinated vulnerabilities, Triage burden, Benchmark contamination, Human-in-the-loop, AI-assisted audit vs AI audit, Prompt injection in auditing pipelines, Responsible disclosure of AI-found bugs, Continuous scanning and run-count coverage, Proof-of-concept harness, Threat model file

Getting help

Firms on this index that handle this in practice: zkSecurity, Trail of Bits, Zellic, Nethermind Security.