Agent Tuning & Optimization 相关度: 9/10

UI-Copilot: Advancing Long-Horizon GUI Automation via Tool-Integrated Policy Optimization

Zhengxi Lu, Fei Tang, Guangyi Liu, Kaitao Song, Xu Tan, Jin Ma, Wenqi Zhang, Weiming Lu, Jun Xiao, Yueting Zhuang, Yongliang Shen
arXiv: 2604.13822v1 发布: 2026-04-15 更新: 2026-04-15

AI 摘要

UI-Copilot通过协同框架和工具集成策略优化,提升了MLLM在长程GUI自动化任务中的性能。

主要贡献

  • 提出了UI-Copilot框架,将GUI Agent和Copilot分离,实现协同任务处理。
  • 引入Memory decoupling,分离持久观测和执行上下文,解决记忆衰退问题。
  • 提出了Tool-Integrated Policy Optimization (TIPO),优化工具选择和任务执行。

方法论

UI-Copilot使用协同框架,通过TIPO分别优化工具选择和任务执行,并采用memory decoupling策略。

原文摘要

MLLM-based GUI agents have demonstrated strong capabilities in complex user interface interaction tasks. However, long-horizon scenarios remain challenging, as these agents are burdened with tasks beyond their intrinsic capabilities, suffering from memory degradation, progress confusion, and math hallucination. To address these challenges, we present UI-Copilot, a collaborative framework where the GUI agent focuses on task execution while a lightweight copilot provides on-demand assistance for memory retrieval and numerical computation. We introduce memory decoupling to separate persistent observations from transient execution context, and train the policy agent to selectively invoke the copilot as Retriever or Calculator based on task demands. To enable effective tool invocation learning, we propose Tool-Integrated Policy Optimization (TIPO), which separately optimizes tool selection through single-turn prediction and task execution through on-policy multi-turn rollouts. Experimental results show that UI-Copilot-7B achieves state-of-the-art performance on challenging MemGUI-Bench, outperforming strong 7B-scale GUI agents such as GUI-Owl-7B and UI-TARS-1.5-7B. Moreover, UI-Copilot-7B delivers a 17.1% absolute improvement on AndroidWorld over the base Qwen model, highlighting UI-Copilot's strong generalization to real-world GUI tasks.

标签

GUI Automation MLLM AI Agent Tool Use Policy Optimization

arXiv 分类

cs.LG