AI Agents 相关度: 9/10

ADAPT: Benchmarking Commonsense Planning under Unspecified Affordance Constraints

Pei-An Chen, Yong-Ching Liang, Jia-Fong Yeh, Hung-Ting Su, Yi-Ting Chen, Min Sun, Winston Hsu
arXiv: 2604.14902v1 发布: 2026-04-16 更新: 2026-04-16

AI 摘要

ADAPT提出DynAfford基准,并引入ADAPT模块,提升具身智能体在动态环境中常识规划的鲁棒性。

主要贡献

  • 提出DynAfford基准,评估具身智能体在动态环境中常识规划能力
  • 引入ADAPT模块,增强现有规划器的可操纵性推理
  • 证明领域自适应的视觉-语言模型在可操纵性推理方面优于GPT-4o

方法论

通过ADAPT模块扩展现有规划器,使其具备显式的可操纵性推理能力。实验对比不同模型在DynAfford上的表现。

原文摘要

Intelligent embodied agents should not simply follow instructions, as real-world environments often involve unexpected conditions and exceptions. However, existing methods usually focus on directly executing instructions, without considering whether the target objects can actually be manipulated, meaning they fail to assess available affordances. To address this limitation, we introduce DynAfford, a benchmark that evaluates embodied agents in dynamic environments where object affordances may change over time and are not specified in the instruction. DynAfford requires agents to perceive object states, infer implicit preconditions, and adapt their actions accordingly. To enable this capability, we introduce ADAPT, a plug-and-play module that augments existing planners with explicit affordance reasoning. Experiments demonstrate that incorporating ADAPT significantly improves robustness and task success across both seen and unseen environments. We also show that a domain-adapted, LoRA-finetuned vision-language model used as the affordance inference backend outperforms a commercial LLM (GPT-4o), highlighting the importance of task-aligned affordance grounding.

标签

具身智能 常识规划 可操纵性推理 动态环境 视觉-语言模型

arXiv 分类

cs.AI cs.CL cs.CV cs.RO