Skip to main content
Aggregate arXiv cs.AI 人工智能 26 Aug 2026 - 13:00

Automata from Agent Traces: Failure and Next-Step Prediction

RSS 官方收录 · 可信分层展示

关键摘要

arXiv:2608.…

  • 23670v1 Announce Type: new Abstract: LLM-based agents execute multi-st…
  • Existing approaches operate per-trace or success-only, so they miss th…
  • To recover that shared structure, we collapse an entire trace corpus i…

摘要引擎:抽取

正文提要

arXiv:2608.23670v1 Announce Type: new Abstract: LLM-based agents execute multi-step tasks, but their behavioral structure remains opaque: long unstructured traces resist the safety auditing and runtime monitoring that deployment requires. Existing approaches operate per-trace or success-only, so they miss the cross-run topology that links next-step and failure prediction. To recover that shared structure, we collapse an entire trace corpus into a single, compact finite-state machine (FSM) that serves as a structural substrate for the otherwise unpredictable behavior of LLM agents. Across twelve public datasets, the FSMs are compact (7-43 states), replay held-out data at >=0.997 fitness with near-identical topology across splits, and build in milliseconds. This substrate addresses both prediction goals. For next-step prediction, FSM-state context outperforms Agent Workflow Memory on every ground-truth-matched dataset. For failure prediction, per-state behavioral features reach held-out AUROC up to 0.94, and an online monitor ranks failing runs above passing ones from a partial trace, triggering early stopping well before completion. Behavioral topology thus appears shaped more by the deployment harness than by the LLM, providing a model-agnostic structural primitive for safety auditing and runtime monitoring.

来源:https://arxiv.org/abs/2608.23670

打开官方原文 站点原文页 可信分区 本信源更多 今日简报 分享图 RSS 稍后再看列表