微信内可能无法直接打开本站。请点右上角 ··· → 在浏览器打开,或复制链接。
FuzzingBrain-Bench V1: Evaluating Open-Ended Bug Discovery by LLMs
RSS 官方收录 · 可信分层展示
关键摘要
arXiv:2608.25158v1 Announce Type: new Abstract: Evaluating the ability of large language models (LLMs) to discover software bugs is increasingly important.…
- Existing benchmarks typically evaluate this capability by asking the m…
- However, this setup may overlook valid crashes discovered by the model…
- As a result, the evaluation may not reflect the model's real capability.
摘要引擎:抽取
正文提要
arXiv:2608.25158v1 Announce Type: new Abstract: Evaluating the ability of large language models (LLMs) to discover software bugs is increasingly important. Existing benchmarks typically evaluate this capability by asking the model to generate a proof-of-concept input that triggers a predefined target vulnerability. However, this setup may overlook valid crashes discovered by the model when they do not match the predefined target. As a result, the evaluation may not reflect the model's real capability. We present FuzzingBrain-Bench, a benchmark for assessing AI models' ability to discover bugs in open-source software. Models are given an open-source project and a sanitizer-instrumented harness in a self-contained Docker image. Their goal is to generate inputs that trigger as many distinct crashes as possible through the harness. A model's performance on each challenge is scored based on the number of distinct crash signatures it produces, capped at a predefined maximum and weighted by a difficulty coefficient. FuzzingBrain-Bench V1 consists of 77 challenges drawn from 43 open-source projects, with 36 C, 32 C++, and 9 Java/JVM challenges. We evaluate Claude Haiku 4.5, Claude Sonnet 4.6, and Claude Opus 4.8 on the full benchmark. Claude Opus 4.8 performs best, triggering crashes in 60 of 77 challenges and achieving a score of 196 out of 579. None of the three models triggers a crash in 13 challenges. The FuzzingBrain-Bench corpus and harnesses are publicly available at https://github.com/fuzzingbrain/FuzzingBrain-Bench.