Skip to main content
Aggregate arXiv cs.AI 人工智能 3 Sep 2026 - 14:00

ToolGate: An Executable Acceptance Pipeline for Tool-Dependent Scientific Benchmark Construction

RSS 官方收录 · 可信分层展示

关键摘要

arXiv:2609.…

  • 02067v1 Announce Type: new Abstract: Scientific benchmarks are commonl…
  • These routes can produce strong evaluations, but they require substant…
  • Language models can reduce this repeated work by proposing candidates …

摘要引擎:抽取

正文提要

arXiv:2609.02067v1 Announce Type: new Abstract: Scientific benchmarks are commonly built by domain experts who write tasks and cross-check one another's work, or who adapt existing material from textbooks, published papers, and online resources. These routes can produce strong evaluations, but they require substantial per-item labor. Language models can reduce this repeated work by proposing candidates quickly. The remaining problem is acceptance. We target scientific questions whose answers require computations with specialist software rather than unaided reasoning alone. A candidate is invalid if its script fails or returns a different answer, or trivial if a model answers it without the software. We present ToolGate, which treats every generated item as a proposal and keeps it only if three gates pass. First, an executable solution script must reproduce the proposed answer when run with the scientific software. Second, randomized no-tool screening rejects candidates that models can already solve from the prompt alone. Third, a tool-using agent must solve each survivor within a fixed time limit. We instantiate ToolGate in FEniCSx with 500 generation attempts. The local-verification gate retains 478 candidates. For final reporting, we rescreen this pool after generation: two randomized no-tool screens exclude 222 from the reported pool, and direct GPT-5.5 API calls at medium reasoning (the API default) exclude another 121. Of the remaining 135, a GPT-5.5 Codex CLI agent with access to FEniCSx solves 130; exact deduplication leaves 128 unique protocol survivors. ToolGate turns repeated answer checking and difficulty screening into an auditable process while leaving domain design and final review to experts.

来源:https://arxiv.org/abs/2609.02067

打开官方原文 站点原文页 可信分区 本信源更多 今日简报 分享图 RSS 稍后再看列表