Multimodal Learning 相关度: 9/10

Towards Unconstrained Human-Object Interaction

Francesco Tonini, Alessandro Conti, Lorenzo Vaquero, Cigdem Beyan, Elisa Ricci
arXiv: 2604.14069v1 发布: 2026-04-15 更新: 2026-04-15

AI 摘要

提出Unconstrained HOI任务,利用MLLM进行无约束人与物交互检测,突破传统HOI的限制。

主要贡献

  • 定义了Unconstrained HOI (U-HOI) 任务
  • 提出了基于MLLM的U-HOI检测pipeline
  • 评估了MLLM在U-HOI上的表现

方法论

利用MLLM直接进行HOI检测,通过测试时推理和语言到图转换提取结构化交互。

原文摘要

Human-Object Interaction (HOI) detection is a longstanding computer vision problem concerned with predicting the interaction between humans and objects. Current HOI models rely on a vocabulary of interactions at training and inference time, limiting their applicability to static environments. With the advent of Multimodal Large Language Models (MLLMs), it has become feasible to explore more flexible paradigms for interaction recognition. In this work, we revisit HOI detection through the lens of MLLMs and apply them to in-the-wild HOI detection. We define the Unconstrained HOI (U-HOI) task, a novel HOI domain that removes the requirement for a predefined list of interactions at both training and inference. We evaluate a range of MLLMs on this setting and introduce a pipeline that includes test-time inference and language-to-graph conversion to extract structured interactions from free-form text. Our findings highlight the limitations of current HOI detectors and the value of MLLMs for U-HOI. Code will be available at https://github.com/francescotonini/anyhoi

标签

HOI Detection Multimodal Learning MLLM Computer Vision Unconstrained Learning

arXiv 分类

cs.CV