safety security
The Bug Test Counted Every Distinct Crash
FuzzingBrain-Bench rewards open-ended crash discovery instead of matching one predefined vulnerability.

Summary
FuzzingBrain-Bench rewards open-ended crash discovery instead of matching one predefined vulnerability.
The first release includes 77 challenges from 43 open-source projects. Of three evaluated models, Claude Opus 4.8 triggered crashes in 60 challenges and scored 196 of 579; none triggered a crash in 13 challenges. Crash signatures measure discovered failures, not exploitability or security severity.
Why it matters
FuzzingBrain-Bench rewards open-ended crash discovery instead of matching one predefined vulnerability.
Limits and context
- Crash signatures measure discovered failures, not exploitability or security severity.
Key claims
FuzzingBrain-Bench rewards open-ended crash discovery instead of matching one predefined vulnerability.
Qualification: Crash signatures measure discovered failures, not exploitability or security severity.
Evidence: source-2026-08-27-013
Sources
- arXiv preprint 2608.25158arXiv · primary research
Corrections
No corrections have been recorded for this story.