Multimodal Learning 相关度: 8/10

MoRight: Motion Control Done Right

Shaowei Liu, Xuanchi Ren, Tianchang Shen, Huan Ling, Saurabh Gupta, Shenlong Wang, Sanja Fidler, Jun Gao
arXiv: 2604.07348v1 发布: 2026-04-08 更新: 2026-04-08

AI 摘要

MoRight提出了一种解耦运动建模框架,用于生成运动控制视频,实现相机视角和物体运动的解耦控制和因果关系建模。

主要贡献

  • 提出解耦的运动建模框架,分离物体运动和相机视角控制
  • 将运动分解为主动和被动成分,学习运动因果关系
  • 实现正向推理和逆向推理,根据主动运动预测结果或根据结果推断动作

方法论

通过时间跨视角注意力将物体运动从规范静态视图转换到任意目标相机视角,并分解运动成分学习因果关系。

原文摘要

Generating motion-controlled videos--where user-specified actions drive physically plausible scene dynamics under freely chosen viewpoints--demands two capabilities: (1) disentangled motion control, allowing users to separately control the object motion and adjust camera viewpoint; and (2) motion causality, ensuring that user-driven actions trigger coherent reactions from other objects rather than merely displacing pixels. Existing methods fall short on both fronts: they entangle camera and object motion into a single tracking signal and treat motion as kinematic displacement without modeling causal relationships between object motion. We introduce MoRight, a unified framework that addresses both limitations through disentangled motion modeling. Object motion is specified in a canonical static-view and transferred to an arbitrary target camera viewpoint via temporal cross-view attention, enabling disentangled camera and object control. We further decompose motion into active (user-driven) and passive (consequence) components, training the model to learn motion causality from data. At inference, users can either supply active motion and MoRight predicts consequences (forward reasoning), or specify desired passive outcomes and MoRight recovers plausible driving actions (inverse reasoning), all while freely adjusting the camera viewpoint. Experiments on three benchmarks demonstrate state-of-the-art performance in generation quality, motion controllability, and interaction awareness.

标签

运动控制 视频生成 因果推理 解耦表示

arXiv 分类

cs.CV cs.AI cs.GR cs.LG cs.RO