Multimodal Learning 相关度: 7/10

Pair2Scene: Learning Local Object Relations for Procedural Scene Generation

Xingjian Ran, Shujie Zhang, Weipeng Zhong, Li Luo, Bo Dai
arXiv: 2604.11808v1 发布: 2026-04-13 更新: 2026-04-13

AI 摘要

Pair2Scene提出了一种基于局部对象关系学习的程序化室内场景生成框架。

主要贡献

  • 提出了Pair2Scene框架,利用局部关系生成场景
  • 构建了3D-Pairs数据集用于训练模型
  • 通过分层结构和碰撞避免实现全局布局

方法论

通过学习对象间的支持和功能关系,使用神经网络估计依赖对象的位置分布,递归生成场景。

原文摘要

Generating high-fidelity 3D indoor scenes remains a significant challenge due to data scarcity and the complexity of modeling intricate spatial relations. Current methods often struggle to scale beyond training distribution to dense scenes or rely on LLMs/VLMs that lack the ability for precise spatial reasoning. Building on top of the observation that object placement relies mainly on local dependencies instead of information-redundant global distributions, in this paper, we propose Pair2Scene, a novel procedural generation framework that integrates learned local rules with scene hierarchies and physics-based algorithms. These rules mainly capture two types of inter-object relations, namely support relations that follow physical hierarchies, and functional relations that reflect semantic links. We model these rules through a network, which estimates spatial position distributions of dependent objects conditioned on position and geometry of the anchor ones. Accordingly, we curate a dataset 3D-Pairs from existing scene data to train the model. During inference, our framework can generate scenes by recursively applying our model within a hierarchical structure, leveraging collision-aware rejection sampling to align local rules into coherent global layouts. Extensive experiments demonstrate that our framework outperforms existing methods in generating complex environments that go beyond training data while maintaining physical and semantic plausibility.

标签

3D Scene Generation Procedural Generation Object Relations Deep Learning

arXiv 分类

cs.CV