SocialMirror: Reconstructing 3D Human Interaction Behaviors from Monocular Videos with Semantic and Geometric Guidance
AI 摘要
SocialMirror利用语义和几何信息,从单目视频重建多人交互的3D人体网格。
主要贡献
- 提出基于扩散模型的框架SocialMirror,解决交互场景下的遮挡问题。
- 利用视觉语言模型指导语义引导的运动填充,恢复遮挡的人体并解决姿态歧义。
- 提出序列级时间细化器,强制平滑运动,并结合几何约束确保合理的接触和空间关系。
方法论
采用基于扩散模型的框架,结合视觉语言模型生成的语义信息和几何约束,进行3D人体网格重建。
原文摘要
Accurately reconstructing human behavior in close-interaction scenarios is crucial for enabling realistic virtual interactions in augmented reality, precise motion analysis in sports, and natural collaborative behavior in human-robot tasks. Reliable reconstruction in these contexts significantly enhances the realism and effectiveness of AI-driven interactive applications. However, human reconstruction from monocular videos in close-interaction scenarios remains challenging due to severe mutual occlusions, leading local motion ambiguity, disrupted temporal continuity and spatial relationship error. In this paper, we propose SocialMirror, a diffusion-based framework that integrates semantic and geometric cues to effectively address these issues. Specifically, we first leverage high-level interaction descriptions generated by a vision-language model to guide a semantic-guided motion infiller, hallucinating occluded bodies and resolving local pose ambiguities. Next, we propose a sequence-level temporal refiner that enforces smooth, jitter-free motions, while incorporating geometric constraints during sampling to ensure plausible contact and spatial relationships. Evaluations on multiple interaction benchmarks show that SocialMirror achieves state-of-the-art performance in reconstructing interactive human meshes, demonstrating strong generalization across unseen datasets and in-the-wild scenarios. The code will be released upon publication.