微信内可能无法直接打开本站。请点右上角 ··· → 在浏览器打开,或复制链接。
Benchmarking General Mobile Assistants in Challenging Real-World Scenarios
RSS 官方收录 · 可信分层展示
关键摘要
arXiv:2608.…
- 27477v1 Announce Type: new Abstract: Graphical user interfaces have em…
- Existing benchmarks such as AndroidWorld and MobileWorld provide stron…
- We present GMA, a benchmark for evaluating general mobile assistants i…
摘要引擎:抽取
正文提要
arXiv:2608.27477v1 Announce Type: new Abstract: Graphical user interfaces have emerged as an important environment for evaluating autonomous AI agents on multimodal interactive tasks. Existing benchmarks such as AndroidWorld and MobileWorld provide strong foundations for mobile agent evaluation, but their application coverage and task design do not yet fully capture the diversity and complexity of realistic mobile use. We present GMA, a benchmark for evaluating general mobile assistants in challenging real-world scenarios. GMA introduces seven applications based on open-source projects, spanning domains such as lifestyle sharing and travel planning, and 300 tasks across four difficulty tiers, from atomic actions to complex multi-step workflows. We evaluate eight frontier models and find that performance declines substantially as task complexity increases, with current agents remaining far from reliably handling realistic user requirements. We further conduct controlled ablation studies of agent harness choices, including context retention and explicit state tracking, under a shared environment, model setting, and task taxonomy. Results show that appropriate harness design can meaningfully improve performance, particularly on demanding workflows, while the effectiveness of specific designs can vary across foundation models. Overall, GMA complements existing benchmarks by expanding application coverage and task complexity, providing a challenging testbed for evaluating mobile agents and studying how harness design supports reliable execution in complex mobile workflows.