Multimodal Learning 相关度: 8/10

3DrawAgent: Teaching LLM to Draw in 3D with Early Contrastive Experience

Hongcan Xiao, Xinyue Xiao, Yilin Wang, Yue Zhang, Yonggang Qi
arXiv: 2604.08042v1 发布: 2026-04-09 更新: 2026-04-09

AI 摘要

3DrawAgent利用LLM和几何反馈生成3D草图,通过对比学习提升3D感知能力,无需参数更新。

主要贡献

  • 提出3DrawAgent框架,用于生成3D草图
  • 使用基于CLIP的感知奖励和LLM评估构建pairwise对比学习
  • 引入relative experience优化策略,提升3D感知能力

方法论

利用LLM驱动,通过几何反馈迭代绘制3D Bezier曲线,并使用pairwise对比学习和relative experience优化策略,实现无参数更新的3D草图生成。

原文摘要

Sketching in 3D space enables expressive reasoning about shape, structure, and spatial relationships, yet generating 3D sketches through natural language remains a major challenge. In this work, we introduce 3DrawAgent, a training-free, language-driven framework for 3D sketch generation that leverages large language models (LLMs) to sequentially draw 3D Bezier curves under geometric feedback. Unlike prior 2D sketch agents, our method introduces a relative experience optimization strategy that adapts the recently proposed Group Reward Policy Optimization (GRPO) paradigm. Instead of relying on explicit ground-truth supervision, we construct pairwise comparisons among generated sketches, with each pair consisting of a relatively better and a worse result based on CLIP-based perceptual rewards and LLM-based fine-grained qualitative assessment. These experiences are then used to iteratively refine the prior knowledge of 3D drawing, enabling black-box reinforcement of the model's 3D awareness. This design allows our model to self-improve its spatial understanding and drawing quality without parameter updates. Experiments show that 3DrawAgent can generate complex and coherent 3D Bezier sketches from diverse textual prompts, exhibit emergent geometric reasoning, and generalize to novel shapes, establishing a new paradigm for advancing the field of training-free 3D sketch intelligence.

标签

3D Sketching Large Language Models Reinforcement Learning Contrastive Learning Geometric Reasoning

arXiv 分类

cs.CV cs.AI