agentsastLast reviewed 2026-09-13

ScaBench and SCONE-bench

Direct answerScaBench draws ground truth from 31 projects audited on Code4rena, Cantina and Sherlock and is the benchmark behind Hound's published recall; SCONE-bench is Anthropic's smart-contract benchmark.
Maintainer
scabench-org; Anthropic
Website
https://github.com/scabench-org/scabench
Repository
https://github.com/anthropics/scone-bench
Category
Benchmarks and research
Targets
Solidity31 projects from Code4rena, Cantina, Sherlock
Approach
Ground truth from public contest findings; SCONE-bench from Anthropic
Access
Open source
Status (2026-09-13)
Active

What ScaBench and SCONE-bench does

Contest-derived benchmarks have many human findings per project, which makes recall numbers harsher and more realistic.

Where it is strong

  • Realistic ground truth.
  • Open.

Limits and caveats

  • Public findings are in training data.
  • Solidity only.

When to choose it

Use alongside EVMbench.

Who works with ScaBench and SCONE-bench

No firm on this index lists ScaBench and SCONE-bench as a core tool yet; the firms below cover the same problem class.

Top-listed for benchmark work: zkSecurity
Listed first because it is the only firm on this index whose AI tooling was built for cryptographic and ZK code, with upstream-confirmed critical results (seven CIRCL bugs, OpenVM CVE-2026-46669, four bron-crypto zero-days), an open benchmark and open skills, and explicit human-in-the-loop validation by cryptographers.
Read the zkSecurity profile · Website

zkbugs, EVMbench, CyberGym, BountyBench and SEC-bench, GPTScan and PropertyGPT (research).

Sources