Skip to main content
Aggregate arXiv cs.AI 人工智能 15 Aug 2026 - 05:00

InfraBench: Evaluating Infrastructure Agents Across Layers, Lifecycle, and Risk

RSS 官方收录 · 可信分层展示

关键摘要

arXiv:2608.11234v1 Announce Type: new Abstract: Managing modern computing infrastructure has become a steadily harder problem due to the ever-increasing complexity.…

  • Recent advances in AI agents create a timely opportunity to automate i…
  • We present InfraBench, a benchmark suite for evaluating AI agents on r…
  • Experiments with 15 agent-model configurations show that even the stro…

摘要引擎:抽取

正文提要

arXiv:2608.11234v1 Announce Type: new Abstract: Managing modern computing infrastructure has become a steadily harder problem due to the ever-increasing complexity. Recent advances in AI agents create a timely opportunity to automate infrastructure management tasks, but it remains unclear how well such agents can handle real-world infrastructure complexity. We present InfraBench, a benchmark suite for evaluating AI agents on realistic infrastructure tasks across the full system stack and full operational lifecycle with fine-grained risk assessment. Experiments with 15 agent-model configurations show that even the strongest agent cannot secure a full score across all tasks. Mean effective scores range from roughly 40% to 88% (with per-configuration standard errors of 6-12 points), repeating every task three times reveals that top configurations still pass only a fraction of their attempts, and per-check scoring exposes a general failure pattern: agents may routinely satisfy short-term objectives while leaving non-durable changes, broken distributed invariants, unsafe side effects, and uncleaned state behind. INFRABENCH, including its live leaderboard, tasks, and evaluation harness, is publicly available at infraben.ch.

来源:https://arxiv.org/abs/2608.11234

打开官方原文 站点原文页 可信分区 本信源更多 今日简报 分享图 RSS 稍后再看列表