Multimodal Learning 相关度: 9/10

What's Missing in Screen-to-Action? Towards a UI-in-the-Loop Paradigm for Multimodal GUI Reasoning

Songze Li, Xiaoke Guo, Tianqi Liu, Biao Yi, Zhaoyan Gong, Zhiqiang Liu, Huajun Chen, Wen Zhang
arXiv: 2604.06995v1 发布: 2026-04-08 更新: 2026-04-08

AI 摘要

提出UILoop范式,通过显式学习UI元素信息,提升MLLM在GUI推理任务中的性能。

主要贡献

  • 提出 UI-in-the-Loop (UILoop) 范式
  • 引入更具挑战性的UI理解任务
  • 构建了包含26K样本的UI Comprehension-Bench基准数据集

方法论

提出将GUI推理视为Screen-UI元素-Action的循环过程,使MLLM学习UI元素定位、语义功能和实际用法。

原文摘要

Existing Graphical User Interface (GUI) reasoning tasks remain challenging, particularly in UI understanding. Current methods typically rely on direct screen-based decision-making, which lacks interpretability and overlooks a comprehensive understanding of UI elements, ultimately leading to task failure. To enhance the understanding and interaction with UIs, we propose an innovative GUI reasoning paradigm called UI-in-the-Loop (UILoop). Our approach treats the GUI reasoning task as a cyclic Screen-UI elements-Action process. By enabling Multimodal Large Language Models (MLLMs) to explicitly learn the localization, semantic functions, and practical usage of key UI elements, UILoop achieves precise element discovery and performs interpretable reasoning. Furthermore, we introduce a more challenging UI Comprehension task centered on UI elements with three evaluation metrics. Correspondingly, we contribute a benchmark of 26K samples (UI Comprehension-Bench) to comprehensively evaluate existing methods' mastery of UI elements. Extensive experiments demonstrate that UILoop achieves state-of-the-art UI understanding performance while yielding superior results in GUI reasoning tasks.

标签

GUI Reasoning Multimodal Learning UI Comprehension Large Language Models

arXiv 分类

cs.AI