Multimodal Learning 相关度: 8/10

UI-Zoomer: Uncertainty-Driven Adaptive Zoom-In for GUI Grounding

Fei Tang, Bofan Chen, Zhengxi Lu, Tongbo Chen, Songqin Nong, Tao Jiang, Wenhao Xu, Weiming Lu, Jun Xiao, Yueting Zhuang, Yongliang Shen
arXiv: 2604.14113v1 发布: 2026-04-15 更新: 2026-04-15

AI 摘要

UI-Zoomer提出了一种无需训练的自适应GUI元素定位缩放框架,提升了小图标和密集布局的定位准确性。

主要贡献

  • 提出了一种基于不确定性驱动的自适应缩放框架UI-Zoomer
  • 利用置信度感知门融合空间共识和token级置信度选择性触发缩放
  • 设计了基于预测方差分解的不确定性驱动的裁剪尺寸模块

方法论

UI-Zoomer通过不确定性量化预测缩放触发和裁剪尺寸,利用置信度感知门和方差分解实现自适应缩放。

原文摘要

GUI grounding, which localizes interface elements from screenshots given natural language queries, remains challenging for small icons and dense layouts. Test-time zoom-in methods improve localization by cropping and re-running inference at higher resolution, but apply cropping uniformly across all instances with fixed crop sizes, ignoring whether the model is actually uncertain on each case. We propose \textbf{UI-Zoomer}, a training-free adaptive zoom-in framework that treats both the trigger and scale of zoom-in as a prediction uncertainty quantification problem. A confidence-aware gate fuses spatial consensus among stochastic candidates with token-level generation confidence to selectively trigger zoom-in only when localization is uncertain. When triggered, an uncertainty-driven crop sizing module decomposes prediction variance into inter-sample positional spread and intra-sample box extent, deriving a per-instance crop radius via the law of total variance. Extensive experiments on ScreenSpot-Pro, UI-Vision, and ScreenSpot-v2 demonstrate consistent improvements over strong baselines across multiple model architectures, achieving gains of up to +13.4\%, +10.3\%, and +4.2\% respectively, with no additional training required.

标签

GUI grounding Uncertainty Quantification Adaptive Zoom-In Computer Vision

arXiv 分类

cs.CV cs.AI cs.CL