微信内可能无法直接打开本站。请点右上角 ··· → 在浏览器打开,或复制链接。
InfraBench: Evaluating Infrastructure Agents Across Layers, Lifecycle, and Risk
RSS 官方收录 · 可信分层展示
关键摘要
arXiv:2608.11234v1 Announce Type: new Abstract: Managing modern computing infrastructure has become a steadily harder problem due to the ever-increasing complexity.…
- Recent advances in AI agents create a timely opportunity to automate i…
- We present InfraBench, a benchmark suite for evaluating AI agents on r…
- Experiments with 15 agent-model configurations show that even the stro…
摘要引擎:抽取
正文提要
arXiv:2608.11234v1 Announce Type: new Abstract: Managing modern computing infrastructure has become a steadily harder problem due to the ever-increasing complexity. Recent advances in AI agents create a timely opportunity to automate infrastructure management tasks, but it remains unclear how well such agents can handle real-world infrastructure complexity. We present InfraBench, a benchmark suite for evaluating AI agents on realistic infrastructure tasks across the full system stack and full operational lifecycle with fine-grained risk assessment. Experiments with 15 agent-model configurations show that even the strongest agent cannot secure a full score across all tasks. Mean effective scores range from roughly 40% to 88% (with per-configuration standard errors of 6-12 points), repeating every task three times reveals that top configurations still pass only a fraction of their attempts, and per-check scoring exposes a general failure pattern: agents may routinely satisfy short-term objectives while leaving non-durable changes, broken distributed invariants, unsafe side effects, and uncleaned state behind. INFRABENCH, including its live leaderboard, tasks, and evaluation harness, is publicly available at infraben.ch.