Skip to main content
Aggregate arXiv cs.AI 人工智能 2 Sep 2026 - 12:30

UI-Venus-2 Technical Report

RSS 官方收录 · 可信分层展示

关键摘要

arXiv:2609.…

  • 00028v1 Announce Type: new Abstract: Multimodal GUI agents have emerge…
  • In this work, we present UI-Venus-2, a general-purpose foundation GUI …
  • To bridge the gap toward practical deployment, we jointly scale three …

摘要引擎:抽取

正文提要

arXiv:2609.00028v1 Announce Type: new Abstract: Multimodal GUI agents have emerged as a promising paradigm for digital task automation, yet transitioning from benchmark-oriented models to dependable real-world applications remains challenging due to limited environment coverage, brittle task construction, and unreliable reward verification. In this work, we present UI-Venus-2, a general-purpose foundation GUI agent designed to operate across mobile, web, and desktop environments through a unified closed-loop reasoning-action framework. To bridge the gap toward practical deployment, we jointly scale three critical dimensions: (1) Environments, expanding coverage to more than 170 multilingual mobile apps and native desktop operating systems; (2) Tasks, employing a deep-research pipeline for function-grounded instruction generation; and (3) Verification, adopting trace-level and sample-level evaluators with visual keypoints and multi-model voting to ensure reliable RL signals for training. Furthermore, we integrate safety-aware mechanisms to ensure controlled execution of consequential actions. By offering a capable, efficient, and open-source foundation, UI-Venus-2 advances the field toward more generalizable, verifiable, and self-reflective agents for real-world applications.

来源:https://arxiv.org/abs/2609.00028

打开官方原文 站点原文页 可信分区 本信源更多 今日简报 分享图 RSS 稍后再看列表