Multimodal Learning 相关度: 8/10

AnyUser: Translating Sketched User Intent into Domestic Robots

Songyuan Yang, Huibin Tan, Kailun Yang, Wenjing Yang, Shaowu Yang
arXiv: 2604.04811v1 发布: 2026-04-06 更新: 2026-04-06

AI 摘要

AnyUser通过草图和语言指令控制机器人,无需预先建模,提升用户交互性和任务完成效率。

主要贡献

  • 提出了一种统一的机器人指令系统AnyUser
  • 实现了多模态输入(草图、视觉、语言)的融合和理解
  • 设计了分层策略用于生成鲁棒的机器人动作

方法论

采用多模态融合理解空间语义信息,结合分层策略生成可执行的机器人动作,并在真实机器人平台上进行验证。

原文摘要

We introduce AnyUser, a unified robotic instruction system for intuitive domestic task instruction via free-form sketches on camera images, optionally with language. AnyUser interprets multimodal inputs (sketch, vision, language) as spatial-semantic primitives to generate executable robot actions requiring no prior maps or models. Novel components include multimodal fusion for understanding and a hierarchical policy for robust action generation. Efficacy is shown via extensive evaluations: (1) Quantitative benchmarks on the large-scale dataset showing high accuracy in interpreting diverse sketch-based commands across various simulated domestic scenes. (2) Real-world validation on two distinct robotic platforms, a statically mounted 7-DoF assistive arm (KUKA LBR iiwa) and a dual-arm mobile manipulator (Realman RMC-AIDAL), performing representative tasks like targeted wiping and area cleaning, confirming the system's ability to ground instructions and execute them reliably in physical environments. (3) A comprehensive user study involving diverse demographics (elderly, simulated non-verbal, low technical literacy) demonstrating significant improvements in usability and task specification efficiency, achieving high task completion rates (85.7%-96.4%) and user satisfaction. AnyUser bridges the gap between advanced robotic capabilities and the need for accessible non-expert interaction, laying the foundation for practical assistive robots adaptable to real-world human environments.

标签

机器人控制 多模态交互 人机交互

arXiv 分类

cs.RO cs.CV cs.HC