AI Agents 相关度: 9/10

ClawGUI: A Unified Framework for Training, Evaluating, and Deploying GUI Agents

Fei Tang, Zhiqiong Lu, Boxuan Zhang, Weiming Lu, Jun Xiao, Yueting Zhuang, Yongliang Shen
arXiv: 2604.11784v1 发布: 2026-04-13 更新: 2026-04-13

AI 摘要

ClawGUI是一个统一的GUI智能体框架,包含训练、评估和部署三部分,并开源了相关工具。

主要贡献

  • 开源GUI智能体RL基础设施
  • 标准化评估流程和基准
  • 支持多种平台和聊天应用的部署

方法论

集成了GiGPO和Process Reward Model进行强化学习,并在统一框架下进行端到端训练。

原文摘要

GUI agents drive applications through their visual interfaces instead of programmatic APIs, interacting with arbitrary software via taps, swipes, and keystrokes, reaching a long tail of applications that CLI-based agents cannot. Yet progress in this area is bottlenecked less by modeling capacity than by the absence of a coherent full-stack infrastructure: online RL training suffers from environment instability and closed pipelines, evaluation protocols drift silently across works, and trained agents rarely reach real users on real devices. We present \textbf{ClawGUI}, an open-source framework addressing these three gaps within a single harness. \textbf{ClawGUI-RL} provides the first open-source GUI agent RL infrastructure with validated support for both parallel virtual environments and real physical devices, integrating GiGPO with a Process Reward Model for dense step-level supervision. \textbf{ClawGUI-Eval} enforces a fully standardized evaluation pipeline across 6 benchmarks and 11+ models, achieving 95.8\% reproduction against official baselines. \textbf{ClawGUI-Agent} brings trained agents to Android, HarmonyOS, and iOS through 12+ chat platforms with hybrid CLI-GUI control and persistent personalized memory. Trained end to end within this pipeline, \textbf{ClawGUI-2B} achieves 17.1\% Success Rate on MobileWorld GUI-Only, outperforming the same-scale MAI-UI-2B baseline by 6.0\%.

标签

GUI Agent Reinforcement Learning Automation Evaluation Deployment

arXiv 分类

cs.LG cs.AI cs.CL cs.CV