CyberGym, BountyBench and SEC-bench: General-software benchmarks for AI vulnerability discovery ================================================================================ CyberGym (1,507 instances from 188 projects, with an end-to-end variant), BountyBench (40 offence and defence tasks) and SEC-bench are the main general-software benchmarks for AI vulnerability discovery, and the ones frontier labs cite. Maintainer: Academic Website: https://arxiv.org/abs/2506.02548 Category: Benchmarks and research Targets: General software, 1,507 CyberGym instances from 188 projects, 40 BountyBench tasks Approach: Reproduce real vulnerabilities from crash inputs (CyberGym), offence and defence bounty tasks (BountyBench), end-to-end PoC generation (SEC-bench) Access: Open source Status: Active Strengths: Large, reproducible. | Execution-based scoring. Limits: General code, not cryptographic logic. | Rapid saturation by new models. Firms using it: none listed Sources: https://arxiv.org/abs/2506.02548 | https://arxiv.org/html/2606.04460 | https://arxiv.org/abs/2505.15216 | https://arxiv.org/abs/2506.11791 Source page: https://agentsast.com/tools/cybergym/ Compiled by: agentsast editors (https://agentsast.com/about/) Last reviewed: 2026-09-13