Multimodal Learning 相关度: 7/10

SocialMirror: Reconstructing 3D Human Interaction Behaviors from Monocular Videos with Semantic and Geometric Guidance

Qi Xia, Peishan Cong, Ziyi Wang, Yujing Sun, Qin Sun, Xinge Zhu, Mao Ye, Ruigang Yang, Yuexin Ma
arXiv: 2604.13581v1 发布: 2026-04-15 更新: 2026-04-15

AI 摘要

SocialMirror利用语义和几何信息,从单目视频重建多人交互的3D人体网格。

主要贡献

  • 提出基于扩散模型的框架SocialMirror,解决交互场景下的遮挡问题。
  • 利用视觉语言模型指导语义引导的运动填充,恢复遮挡的人体并解决姿态歧义。
  • 提出序列级时间细化器,强制平滑运动,并结合几何约束确保合理的接触和空间关系。

方法论

采用基于扩散模型的框架,结合视觉语言模型生成的语义信息和几何约束,进行3D人体网格重建。

原文摘要

Accurately reconstructing human behavior in close-interaction scenarios is crucial for enabling realistic virtual interactions in augmented reality, precise motion analysis in sports, and natural collaborative behavior in human-robot tasks. Reliable reconstruction in these contexts significantly enhances the realism and effectiveness of AI-driven interactive applications. However, human reconstruction from monocular videos in close-interaction scenarios remains challenging due to severe mutual occlusions, leading local motion ambiguity, disrupted temporal continuity and spatial relationship error. In this paper, we propose SocialMirror, a diffusion-based framework that integrates semantic and geometric cues to effectively address these issues. Specifically, we first leverage high-level interaction descriptions generated by a vision-language model to guide a semantic-guided motion infiller, hallucinating occluded bodies and resolving local pose ambiguities. Next, we propose a sequence-level temporal refiner that enforces smooth, jitter-free motions, while incorporating geometric constraints during sampling to ensure plausible contact and spatial relationships. Evaluations on multiple interaction benchmarks show that SocialMirror achieves state-of-the-art performance in reconstructing interactive human meshes, demonstrating strong generalization across unseen datasets and in-the-wild scenarios. The code will be released upon publication.

标签

3D人体重建 单目视频 人机交互 扩散模型 视觉语言模型

arXiv 分类

cs.CV