agentsastLast reviewed 2026-09-13

All tools, by category

Direct answerEvery tool on this index in one place, grouped by what it targets: cryptography and ZK specialists, frontier-lab general scanners, the DARPA AIxCC cyber reasoning systems, smart-contract auditors, and the benchmarks used to measure them. Each row links to a page with approach, access model, published results and limits.

Cryptography and ZK specialists

Tools built for the code that general scanners handle worst: finite-field arithmetic, constraint systems, pairing libraries, MPC and post-quantum implementations. zkao (zkSecurity) is the only product in this class with public, upstream-confirmed critical findings; zk-skills makes the same audit patterns available as open-source agent skills; zkCraft adds LLM guidance to circuit fuzzing. Category guide →

ToolTargetsApproachAccessStatus
zkao
zkSecurity
CircomLeo (Aleo)Rust cryptoGo cryptoMPCFHEPost-quantumTLS / E2EEMulti-agent LLM workflows with expert-maintained skills, a second wave of validating agents, PoC harnesses, and automatic re-scans as models improveSaaS; prepaid non-expiring credits; enterprise plans with human auditsActive (zkao 2.0 released 2026-07-24)
zk-skills and circom-auditor
zkSecurity
CircomClaude CodeCodexCursorAgent skills (prompts, workflows and checklists) that turn a general coding agent into a Circom auditorOpen source (MIT)Active (released 2026-08-05)
zkCraft (with zkFuzz)
Academic (Takahashi et al.)
CircomNoir (preliminary)LLM proposes mutation patterns; zkFuzz's trace-constraint consistency test supplies ground truthOpen sourceResearch (zkFuzz at IEEE S&P 2026; zkCraft 2026 preprint)
AI Grinding for cryptanalysis (research)
Olejnik and Naskrecki (academic)
Published cryptographic constructionsCryptanalysisAutonomous workflow: agents propose low-precision attack hypotheses, an exact, adversarially controlled test provides evidenceResearch paperResearch (2026-08-22)

Frontier-lab and general scanners

General-purpose vulnerability scanners from Anthropic, OpenAI and Google, plus independent products such as AISLE and XBOW. They target C, C++, and mainstream application code and have produced CVEs in OpenSSL, OpenSSH, GnuTLS, wolfSSL, SQLite, FFmpeg and V8. They are not cryptography-aware, but cryptographic libraries are written in the languages they scan. Category guide →

ToolTargetsApproachAccessStatus
Claude Security
Anthropic
General codeEnterprise repositoriesClaude Code pluginAgentic multi-stage analysis that traces data flows and re-examines findings to filter false positives; findings carry CWE, severity and confidence; produces patch filesEnterprise SaaS, billed as token usageActive (public beta May 2026; on Claude Mythos 5 from 2026-08-21)
Codex Security (formerly Aardvark)
OpenAI
General codeCommits and pull requestsBuilds a project threat model, scans commits, validates exploitability in a sandbox, proposes patchesSaaS for ChatGPT Pro, Business, Enterprise and EduActive (research preview 2026-03-06)
Big Sleep and CodeMender
Google DeepMind and Project Zero
C / C++ open sourceV8SQLiteFFmpegLLM agent evolved from Project Naptime; CodeMender validates with sandboxed PoCs and patches with a model-as-judge; Gemini 3.5 Flash Cyber trained on OSV and OSS-Fuzz dataBig Sleep internal; CodeMender preview on Google Cloud; Flash Cyber gated to governments and partnersActive
AISLE
AISLE
C sourceOpenSSLcurlAutonomous analysis of C code with on-premises deployment optionEnterpriseActive
XBOW
XBOW
Web applicationsDeployed servicesAutonomous black-box penetration testing agentSaaSActive (155 million dollar Series C in 2026)

Cyber reasoning systems (DARPA AIxCC)

The seven finalists of DARPA's AI Cyber Challenge, all open-sourced after the August 2025 final. They combine LLMs with fuzzing and program analysis to find and patch bugs in C and Java, processed 54 million lines of code in the final, found 18 real zero-days and patched 43 of 54 synthetic bugs. OpenSSF's OSS-CRS packages them for open-source maintainers. Category guide →

ToolTargetsApproachAccessStatus
Atlantis
Team Atlanta (Georgia Tech, Samsung Research, KAIST, POSTECH)
CJavaEnsemble of independent bug-finding modules sharing seeds, with eight patching agentsOpen sourceOpen-sourced after the 2025-08-08 final
Buttercup
Trail of Bits
CJavaLLM plus fuzzing plus program analysis using non-reasoning models; finds and patchesOpen sourceOpen-sourced 2025
RoboDuck
Theori
CJavaLLM-only pipeline with no fuzzing or symbolic executionOpen sourceOpen-sourced 2025
OSS-CRS and other AIxCC finalists
OpenSSF and the AIxCC finalist teams
CJavaOSS-Fuzz projectsPackaging of finalist components (Shellphish ARTIPHISHELL, 42-b3yond-6ug BugBuster, all-you-need-is-a-fuzzing-brain, Lacrosse) for open-source maintainersOpen sourceActive

Smart-contract AI auditors

Commercial and open-source AI auditors for Solidity, Vyper, Rust (Solana), Move and Cairo. Published recall against human audits ranges from about 30 percent (Nethermind AuditAgent on its own audits) to about 70 percent on the EVMbench benchmark; precision on live code is around 55 percent in the one controlled study (Sherlock AI). Several also cover ZK circuit languages. Category guide →

ToolTargetsApproachAccessStatus
Sherlock AI
Sherlock
SolidityEVMMulti-step LLM reasoning trained on top researchers' findings; GitHub PR integration; codebase chat; verification tests for fixesCommercial (contact sales)Active (v2 May 2026)
AuditAgent
Nethermind Security
EVMSolanaStarknetLLM agent run after manual review as a second layerSaaSActive
Zellic V12
Zellic
SolidityLLM combined with static analysis, aimed at the roughly 70 percent of bugs that are coding mistakesAnnounced as free; current availability and pricing not confirmedActive (announced 2025-09-25)
Savant Chat
Novel Codes DMCC
SolidityVyperRustMoveCairoFunCCircomHalo2NoirarkworksMulti-agent LLM stack across 200+ vulnerability classes; critic subagent writes a PoC per finding on higher tiersPay per line (0.07 to 0.50 dollars) or 250 to 2,500 dollars per month; 75 dollars free creditsActive
Olympix
Olympix
SolidityIR and custom detectors, symbolic execution, fuzzing, mutation testing plus AI, with executable PoCs, run per commitCommercial CI toolActive (founded 2022)
Octane Security
Octane
EVMSolanaAptosSuiCosmosContinuous AI scanning with automated fixesCommercialActive (6.75 million dollar seed)
Hound
Bernhard Mueller (scabench-org)
Language-agnosticSolidityRelation-first knowledge graphs, persistent vulnerability hypotheses, scout and strategist model switchingOpen sourceActive (paper 2025-10)
QuillShield
QuillAudits
SolidityAI audits plus open-source Claude skills using a 'Semantic State Protocol' (behavioral decomposition, threat modeling, adversarial simulation, risk scoring)Commercial; skills open sourceActive
Cecuro
Cecuro
DeFi contractsSpecialised agent; benchmark and baseline open-sourced, agent withheldCommercialActive
Certora AI Composer
Certora
SoliditySecure generation: model writes code, the formal prover checks invariants before acceptanceOpen source alpha (2025-12-04)Alpha
Immunefi Magnus
Immunefi
Smart contractsBounty programsSecurity Swarm agents, Fuzzland AI fuzzing integration, CODEX vulnerability datasetCommercial platformActive

Benchmarks and research

The datasets used to score AI bug finders, and the research systems that established the methods. zkbugs is the only ZK-specific benchmark; EVMbench (OpenAI and Paradigm) is the most cited for Solidity and the most criticised for contamination; CyberGym and BountyBench cover general software. Read the methodology before the headline number. Category guide →

ToolTargetsApproachAccessStatus
zkbugs
zkSecurity
CircomZK DSLs139 catalogued vulnerabilitiesReproducible vulnerable circuits with direct and full-codebase evaluation modes; public knowledge base at bugs.zksecurity.xyzOpen sourceActive
EVMbench
OpenAI and Paradigm
Solidity117 vulnerabilities from 40 auditsDetect, patch and exploit modesOpen sourceActive (released 2026-02-18)
ScaBench and SCONE-bench
scabench-org; Anthropic
Solidity31 projects from Code4rena, Cantina, SherlockGround truth from public contest findings; SCONE-bench from AnthropicOpen sourceActive
CyberGym, BountyBench and SEC-bench
Academic
General software1,507 CyberGym instances from 188 projects40 BountyBench tasksReproduce real vulnerabilities from crash inputs (CyberGym), offence and defence bounty tasks (BountyBench), end-to-end PoC generation (SEC-bench)Open sourceActive
GPTScan and PropertyGPT (research)
Academic
SolidityGPTScan: GPT plus static analysis for logic bugs (ICSE 2024); PropertyGPT: retrieval-augmented generation of formal properties (NDSS 2025)ResearchPublished