Multimodal Learning 相关度: 9/10

LMMs Meet Object-Centric Vision: Understanding, Segmentation, Editing and Generation

Yuqian Yuan, Wenqiao Zhang, Juekai Lin, Yu Zhong, Mingjian Gao, Binhe Yu, Yunqi Cao, Wentong Li, Yueting Zhuang, Beng Chin Ooi
arXiv: 2604.11789v1 发布: 2026-04-13 更新: 2026-04-13

AI 摘要

综述LMM在物体中心视觉方面的进展,涉及理解、分割、编辑和生成。

主要贡献

  • 总结LMMs在物体中心视觉方面的最新进展
  • 将文献组织为四个主题:理解、分割、编辑和生成
  • 讨论了开放性挑战和未来方向

方法论

综述并组织相关文献,从建模范式、学习策略和评估协议等方面进行分析。

原文摘要

Large Multimodal Models (LMMs) have achieved remarkable progress in general-purpose vision--language understanding, yet they remain limited in tasks requiring precise object-level grounding, fine-grained spatial reasoning, and controllable visual manipulation. In particular, existing systems often struggle to identify the correct instance, preserve object identity across interactions, and localize or modify designated regions with high precision. Object-centric vision provides a principled framework for addressing these challenges by promoting explicit representations and operations over visual entities, thereby extending multimodal systems from global scene understanding to object-level understanding, segmentation, editing, and generation. This paper presents a comprehensive review of recent advances at the convergence of LMMs and object-centric vision. We organize the literature into four major themes: object-centric visual understanding, object-centric referring segmentation, object-centric visual editing, and object-centric visual generation. We further summarize the key modeling paradigms, learning strategies, and evaluation protocols that support these capabilities. Finally, we discuss open challenges and future directions, including robust instance permanence, fine-grained spatial control, consistent multi-step interaction, unified cross-task modeling, and reliable benchmarking under distribution shift. We hope this paper provides a structured perspective on the development of scalable, precise, and trustworthy object-centric multimodal systems.

标签

LMM Object-Centric Vision Multimodal Learning

arXiv 分类

cs.CV