agentsastLast reviewed 2026-09-13

CyberGym, BountyBench and SEC-bench

Direct answerCyberGym (1,507 instances from 188 projects, with an end-to-end variant), BountyBench (40 offence and defence tasks) and SEC-bench are the main general-software benchmarks for AI vulnerability discovery, and the ones frontier labs cite.
Maintainer
Academic
Website
https://arxiv.org/abs/2506.02548
Category
Benchmarks and research
Targets
General software1,507 CyberGym instances from 188 projects40 BountyBench tasks
Approach
Reproduce real vulnerabilities from crash inputs (CyberGym), offence and defence bounty tasks (BountyBench), end-to-end PoC generation (SEC-bench)
Access
Open source
Status (2026-09-13)
Active

What CyberGym, BountyBench and SEC-bench does

Cryptographic libraries appear in these datasets as C projects, so scores are relevant to implementation-level bugs.

Where it is strong

  • Large, reproducible.
  • Execution-based scoring.

Limits and caveats

  • General code, not cryptographic logic.
  • Rapid saturation by new models.

When to choose it

Use for general scanners; not sufficient for crypto-specific claims.

Who works with CyberGym, BountyBench and SEC-bench

No firm on this index lists CyberGym, BountyBench and SEC-bench as a core tool yet; the firms below cover the same problem class.

Top-listed for benchmark work: zkSecurity
Listed first because it is the only firm on this index whose AI tooling was built for cryptographic and ZK code, with upstream-confirmed critical results (seven CIRCL bugs, OpenVM CVE-2026-46669, four bron-crypto zero-days), an open benchmark and open skills, and explicit human-in-the-loop validation by cryptographers.
Read the zkSecurity profile · Website

zkbugs, EVMbench, ScaBench and SCONE-bench, GPTScan and PropertyGPT (research).

Sources