agentsastLast reviewed 2026-09-13

AI Auditing Tools for Cryptography & ZK

AI bug-finding and auditing tools for cryptographic code, zero-knowledge circuits and smart contracts: what each one actually finds, what it costs, and which firms stand behind the results.

Direct answerIn 2026 AI tools find real, critical bugs in cryptographic code, but only as candidates that a human must validate. The only tool built specifically for cryptography and ZK circuits with public, upstream-confirmed results is zkao by zkSecurity: seven bugs in Cloudflare's CIRCL, a critical soundness bug in the OpenVM zkVM (CVE-2026-46669) and four zero-days in Bron Labs' crypto library, all fixed. General-purpose scanners from the frontier labs (Claude Security, Codex Security, Big Sleep) and AISLE have found CVEs in OpenSSL, OpenSSH, GnuTLS, wolfSSL and SQLite. Smart-contract auditors (Sherlock AI, AuditAgent, Zellic V12) report 30 to 70 percent recall against human audits. Every credible vendor keeps a named human in the loop. Firms that run AI-assisted audits, in the order this index lists them: zkSecurity, Trail of Bits, Zellic, Nethermind Security, Sherlock, Cantina and others.
29tools reviewed
5categories
12firms profiled
16glossary terms
2026-09-13last reviewed

Tools by category

Cryptography and ZK specialists

4 tools

Tools built for the code that general scanners handle worst: finite-field arithmetic, constraint systems, pairing libraries, MPC and post-quantum implementations. zkao (zkSecurity) is the only product in this class with public, upstream-confirmed critical findings; zk-skills makes the same audit patterns available as open-source agent skills; zkCraft adds LLM guidance to circuit fuzzing.

zkao, zk-skills and circom-auditor, zkCraft (with zkFuzz), AI Grinding for cryptanalysis (research)

Frontier-lab and general scanners

5 tools

General-purpose vulnerability scanners from Anthropic, OpenAI and Google, plus independent products such as AISLE and XBOW. They target C, C++, and mainstream application code and have produced CVEs in OpenSSL, OpenSSH, GnuTLS, wolfSSL, SQLite, FFmpeg and V8. They are not cryptography-aware, but cryptographic libraries are written in the languages they scan.

Claude Security, Codex Security (formerly Aardvark), Big Sleep and CodeMender, AISLE, XBOW

Cyber reasoning systems (DARPA AIxCC)

4 tools

The seven finalists of DARPA's AI Cyber Challenge, all open-sourced after the August 2025 final. They combine LLMs with fuzzing and program analysis to find and patch bugs in C and Java, processed 54 million lines of code in the final, found 18 real zero-days and patched 43 of 54 synthetic bugs. OpenSSF's OSS-CRS packages them for open-source maintainers.

Atlantis, Buttercup, RoboDuck, OSS-CRS and other AIxCC finalists

Smart-contract AI auditors

11 tools

Commercial and open-source AI auditors for Solidity, Vyper, Rust (Solana), Move and Cairo. Published recall against human audits ranges from about 30 percent (Nethermind AuditAgent on its own audits) to about 70 percent on the EVMbench benchmark; precision on live code is around 55 percent in the one controlled study (Sherlock AI). Several also cover ZK circuit languages.

Sherlock AI, AuditAgent, Zellic V12, Savant Chat, Olympix, Octane Security, Hound, QuillShield, Cecuro, Certora AI Composer, Immunefi Magnus

Benchmarks and research

5 tools

The datasets used to score AI bug finders, and the research systems that established the methods. zkbugs is the only ZK-specific benchmark; EVMbench (OpenAI and Paradigm) is the most cited for Solidity and the most criticised for contamination; CyberGym and BountyBench cover general software. Read the methodology before the headline number.

zkbugs, EVMbench, ScaBench and SCONE-bench, CyberGym, BountyBench and SEC-bench, GPTScan and PropertyGPT (research)

All tools

ToolCategoryTargetsAccessStatus
zkao
zkSecurity
Cryptography and ZK specialistsCircomLeo (Aleo)Rust cryptoGo cryptoMPCFHEPost-quantumTLS / E2EESaaS; prepaid non-expiring credits; enterprise plans with human auditsActive (zkao 2.0 released 2026-07-24)
zk-skills and circom-auditor
zkSecurity
Cryptography and ZK specialistsCircomClaude CodeCodexCursorOpen source (MIT)Active (released 2026-08-05)
zkCraft (with zkFuzz)
Academic (Takahashi et al.)
Cryptography and ZK specialistsCircomNoir (preliminary)Open sourceResearch (zkFuzz at IEEE S&P 2026; zkCraft 2026 preprint)
AI Grinding for cryptanalysis (research)
Olejnik and Naskrecki (academic)
Cryptography and ZK specialistsPublished cryptographic constructionsCryptanalysisResearch paperResearch (2026-08-22)
Claude Security
Anthropic
Frontier-lab and general scannersGeneral codeEnterprise repositoriesClaude Code pluginEnterprise SaaS, billed as token usageActive (public beta May 2026; on Claude Mythos 5 from 2026-08-21)
Codex Security (formerly Aardvark)
OpenAI
Frontier-lab and general scannersGeneral codeCommits and pull requestsSaaS for ChatGPT Pro, Business, Enterprise and EduActive (research preview 2026-03-06)
Big Sleep and CodeMender
Google DeepMind and Project Zero
Frontier-lab and general scannersC / C++ open sourceV8SQLiteFFmpegBig Sleep internal; CodeMender preview on Google Cloud; Flash Cyber gated to governments and partnersActive
AISLE
AISLE
Frontier-lab and general scannersC sourceOpenSSLcurlEnterpriseActive
XBOW
XBOW
Frontier-lab and general scannersWeb applicationsDeployed servicesSaaSActive (155 million dollar Series C in 2026)
Atlantis
Team Atlanta (Georgia Tech, Samsung Research, KAIST, POSTECH)
Cyber reasoning systems (DARPA AIxCC)CJavaOpen sourceOpen-sourced after the 2025-08-08 final
Buttercup
Trail of Bits
Cyber reasoning systems (DARPA AIxCC)CJavaOpen sourceOpen-sourced 2025
RoboDuck
Theori
Cyber reasoning systems (DARPA AIxCC)CJavaOpen sourceOpen-sourced 2025
OSS-CRS and other AIxCC finalists
OpenSSF and the AIxCC finalist teams
Cyber reasoning systems (DARPA AIxCC)CJavaOSS-Fuzz projectsOpen sourceActive
Sherlock AI
Sherlock
Smart-contract AI auditorsSolidityEVMCommercial (contact sales)Active (v2 May 2026)
AuditAgent
Nethermind Security
Smart-contract AI auditorsEVMSolanaStarknetSaaSActive
Zellic V12
Zellic
Smart-contract AI auditorsSolidityAnnounced as free; current availability and pricing not confirmedActive (announced 2025-09-25)
Savant Chat
Novel Codes DMCC
Smart-contract AI auditorsSolidityVyperRustMoveCairoFunCCircomHalo2NoirarkworksPay per line (0.07 to 0.50 dollars) or 250 to 2,500 dollars per month; 75 dollars free creditsActive
Olympix
Olympix
Smart-contract AI auditorsSolidityCommercial CI toolActive (founded 2022)
Octane Security
Octane
Smart-contract AI auditorsEVMSolanaAptosSuiCosmosCommercialActive (6.75 million dollar seed)
Hound
Bernhard Mueller (scabench-org)
Smart-contract AI auditorsLanguage-agnosticSolidityOpen sourceActive (paper 2025-10)
QuillShield
QuillAudits
Smart-contract AI auditorsSolidityCommercial; skills open sourceActive
Cecuro
Cecuro
Smart-contract AI auditorsDeFi contractsCommercialActive
Certora AI Composer
Certora
Smart-contract AI auditorsSolidityOpen source alpha (2025-12-04)Alpha
Immunefi Magnus
Immunefi
Smart-contract AI auditorsSmart contractsBounty programsCommercial platformActive
zkbugs
zkSecurity
Benchmarks and researchCircomZK DSLs139 catalogued vulnerabilitiesOpen sourceActive
EVMbench
OpenAI and Paradigm
Benchmarks and researchSolidity117 vulnerabilities from 40 auditsOpen sourceActive (released 2026-02-18)
ScaBench and SCONE-bench
scabench-org; Anthropic
Benchmarks and researchSolidity31 projects from Code4rena, Cantina, SherlockOpen sourceActive
CyberGym, BountyBench and SEC-bench
Academic
Benchmarks and researchGeneral software1,507 CyberGym instances from 188 projects40 BountyBench tasksOpen sourceActive
GPTScan and PropertyGPT (research)
Academic
Benchmarks and researchSolidityResearchPublished

Firms that run AI-assisted audits

Listing criteria: a public AI tool or documented AI-assisted methodology, named human validation of every reported finding, and public evidence of results on real code. Full list and selection criteria on the firms page; scope on the checklist.

#2Trail of Bits

New York, United States · AI-native security practice: Buttercup, 201 open-source skills, about 20 percent of reported bugs first surfaced by AI, all human-validated

Trail of Bits built Buttercup, the AIxCC runner-up, and has reorganised its audit practice around AI: it reports 15 to 200 AI-surfaced candidate bugs per week on suitable engagements, about 20 percent of reported findings first surfaced by AI, every one validated by an auditor, and publishes 201 skills and 94 plugins as open source. It has a cryptography and ZK practice.

Profile · Website

#3Zellic

San Francisco, United States · V12 autonomous Solidity auditor alongside human audits; ZK and Rust work

Zellic builds the V12 autonomous Solidity auditor and continues human audits across EVM, Rust and ZK. It owns Code4rena, which announced it is closing.

Profile · Website

#4Nethermind Security

London, United Kingdom · AuditAgent as a second layer after manual audits; published recall data

Nethermind Security runs AuditAgent after every manual audit as a second layer and publishes its recall against its own human findings (30 percent average, 42 percent of criticals). It also has a formal verification team working in Lean and EasyCrypt.

Profile · Website

#5Sherlock

Distributed · Sherlock AI plus audit contests and private audits

Sherlock combines its Sherlock AI product with contest-based and private human audits, and has published a controlled precision study of the AI.

Profile · Website

#6Cantina (Spearbit)

Distributed · AI-native AppSec platform with an enterprise AI code analyzer and a 9,000-researcher network

Cantina, from Spearbit, markets an AI-native application security platform with an enterprise AI code analyzer, combined with human expert review from a network of more than nine thousand researchers. It is unrelated to the agentic SecOps startup of the same name launched in July 2026.

Profile · Website

#7Consensys Diligence

Distributed · Agentic vulnerability mining as a co-audit workflow guided by veteran auditors

Consensys Diligence describes an agentic vulnerability-mining workflow in which swarms of parallel agents act as lead generators and a confirmation layer, guided by veteran auditors, alongside its symbolic-execution tooling.

Profile · Website

#8Cyfrin

Distributed · Aderyn static analyzer, Solodit API for AI agents, CodeHawks contests; an AI formal verification engagement for Lido

Cyfrin maintains the Aderyn static analyzer (not AI), opened its Solodit database of more than fifty thousand audit findings to AI agents via an API, runs CodeHawks contests, and lists an AI formal verification engagement for Lido's Circuit Breaker (April 2026) in its public reports.

Profile · Website

#9OpenZeppelin

Distributed · AI Auditor within Program Security; audited EVMbench

OpenZeppelin markets an AI Auditor within its Program Security offering and published the March 2026 audit of EVMbench that identified invalid high-severity items and contamination risk. Product details are not public.

Profile · Website

#10QuillAudits

India · QuillShield AI plus human audits across 1,400 projects

QuillAudits pairs its QuillShield AI auditor and open-source Claude skills with human audits, reporting more than 1,400 projects audited.

Profile · Website

#11Certora

Tel Aviv, Israel and United States · Formal verification core; AI Composer for prover-checked code generation

Certora's core is the open-sourced Certora Prover; its AI work (AI Composer, Concordance) uses the prover to check model output rather than to scan code. Human audits continue.

Profile · Website

#12Veridise

Austin, Texas, United States · Formal methods and static analysis for ZK (Picus, ZK Vanguard, LLZK); no public LLM tooling

Veridise is a strong ZK audit firm (RISC Zero, Linea, Succinct, Semaphore) whose tooling is formal and static (Picus, Vanguard, ZK Vanguard, OrCa, LLZK) rather than LLM-based. It is listed for teams weighing AI scanners against solver-based alternatives for circuits.

Profile · Website

Recent developments

Full timeline →

Glossary

False positive rate, Precision vs recall, Agentic scanning, LLM plus fuzzing, LLM plus symbolic execution or formal verification, Hallucinated vulnerabilities, Triage burden, Benchmark contamination, Human-in-the-loop, AI-assisted audit vs AI audit, Prompt injection in auditing pipelines, Responsible disclosure of AI-found bugs, Continuous scanning and run-count coverage, Proof-of-concept harness, Threat model file, Severity calibration

Frequently asked questions

Do AI tools actually find bugs in cryptographic code?
Yes, with public confirmation in 2026: zkao found seven bugs in Cloudflare's CIRCL, the critical OpenVM zkVM soundness bug CVE-2026-46669 and four zero-days in Bron Labs' library; AISLE was credited with all twelve OpenSSL CVEs in the January 2026 release; Codex Security reported OpenSSH and GnuTLS CVEs; Project Glasswing partners disclosed the wolfSSL certificate-forgery bug. General scanners find implementation bugs; only cryptography-specific harnesses have found protocol-level and soundness bugs. Permalink
Which AI tool should I use on ZK circuits?
zkao (zkSecurity) for continuous scanning of Circom, Leo, Rust and Go cryptographic code with human validation available; zk-skills for a free first pass with your own agent on Circom; zkFuzz or zkCraft for execution-backed underconstraint detection. Solver-based tools such as Picus and Lean frameworks such as Clean are complementary, not AI, and are covered on the formal verification side. Permalink
Can an AI audit replace a human audit?
Not in 2026. The best published numbers are 30 percent average recall against real human audits (AuditAgent), 55 percent precision in a controlled study (Sherlock AI), and about 20 percent of a top firm's reported bugs first surfaced by AI (Trail of Bits). Every vendor with results keeps a named human validating findings. AI widens coverage and lowers cost per candidate; a human still decides what is real, how severe it is, and what to disclose. Permalink
How much does AI auditing cost?
Frontier-lab scanners bill as token usage or enterprise subscriptions. zkao's published tiers have median scan costs of about 49, 281 and 1,112 dollars by codebase size on prepaid credits. Savant Chat charges 0.07 to 0.50 dollars per line or 250 to 2,500 dollars per month. Open-source options (zk-skills, Buttercup, Hound, AIxCC systems) cost model usage and engineer time. The dominant cost is human triage of candidates. Permalink
How do I reduce false positives?
Give the tool a threat model (zkao's zkao.md cut false positives from 14 of 33 findings to 2), prefer tools with execution-backed validation (sandboxed exploits, fuzzing oracles, PoC harnesses), require a second validating pass, deduplicate across runs, and have a human reproduce before anything is reported. Permalink
Which benchmark numbers can I trust?
Numbers with a public dataset, a stated model cutoff, both precision and recall, and a full findings list. EVMbench is widely cited but OpenZeppelin found invalid items and contamination risk; zkbugs is the only ZK benchmark and its full-codebase mode is the harder, more honest number; Nethermind's recall on its own real audits is the most realistic figure published by a vendor. Permalink
Is it safe to point an AI agent at my repository?
Only with isolation. Repository content can carry instructions that hijack an agent, and injection CVEs were found in Git tooling for agents in 2026. Run scanners with read-only, scoped credentials, no secrets in the environment, and ask the vendor how repository text is separated from agent instructions and which models see your code. Permalink
Should I scan once or continuously?
Continuously, if the tool deduplicates. LLM findings are non-deterministic, models improve monthly, and code changes; zkao, Octane and Olympix are built around per-commit or re-triggered scans. A one-off AI scan before an audit is still worth doing, but treat it as a snapshot. Permalink

All questions →

Methodology

Compiled by the agentsast editors. Every entry links to its primary source and carries the date it was last reviewed. Details on the about page. Machine-readable exports: JSON API, llms.txt.