Skip to main content
Aggregate AI 摘要 arXiv cs.AI 人工智能 7 Sep 2026 - 15:00

ERPBench: Evaluating LLM Agents for Enterprise Decision-Making Across Competitive Market Ecologies

RSS 官方收录 · 可信分层展示

关键摘要

ERPBench测试100个企业决策问题,揭示LLM在Solo与Arena两种竞争生态中排名差异显著

  • ERPBench含六轮ERP仿真,覆盖定价、生产、采购等六大模块
  • 在Solo生态中DeepSeek估值最高(252.29M),Arena生态中Gemini领先(263.95M)
  • 两生态仅21/100问题选出相同任务级优胜模型,Gemini底排率从22%降至0%

AI 摘要 · 来源可核验

正文提要

arXiv:2609.04667v1 Announce Type: new Abstract: Large language model (LLM) agents are increasingly proposed for enterprise workflows, yet existing evaluations rarely test whether business-decision conclusions transfer across competitive market ecologies. We introduce ERPBench, an execution-instrumented benchmark for enterprise decision agents in a six-round Enterprise Resource Planning (ERP) simulation with coupled pricing, production, procurement, inventory, finance, and shared-market competition. ERPBench evaluates the same 100 fixed problems in two matched competitive market ecologies: Solo, where each evaluated LLM agent competes against fixed rule-based opponents, and Arena, where six evaluated LLM agents compete in a shared market. Across six model families, this yields 1,200 model-level trajectories spanning 7,200 decision rounds. Under the observed service configuration, the leading model differs between ecologies: DeepSeek leads in Solo (252.29M mean valuation; mean rank 1.67), whereas Gemini leads in Arena (263.95M; 1.76). The two ecologies identify the same task-level winner on only 21 of 100 problems, and Gemini's bottom-rank rate falls from 22 % to 0 % in Arena. ERPBench supports paired evaluation of whether enterprise-agent rankings transfer across competitive market ecologies, supplemented by aggregate execution-intervention analysis. Code and benchmark resources are available in our https://github.com/GAIR-NLP/erp-bench.

来源:https://arxiv.org/abs/2609.04667

打开官方原文 站点原文页 可信分区 本信源更多 今日简报 分享图 RSS 稍后再看列表