Multimodal Learning 相关度: 9/10

Geoparsing: Diagram Parsing for Plane and Solid Geometry with a Unified Formal Language

Peijie Wang, Ming-Liang Zhang, Jun Cao, Chao Deng, Dekang Ran, Hongda Sun, Pi Bu, Xuan Zhang, Yingyao Wang, Jun Song, Bo Zheng, Fei Yin, Cheng-Lin Liu
arXiv: 2604.11600v1 发布: 2026-04-13 更新: 2026-04-13

AI 摘要

该论文提出了一种统一的几何形式化语言用于解析平面和立体几何,并构建了大规模数据集GDP-29K。

主要贡献

  • 设计统一的平面和立体几何形式化语言
  • 构建大规模几何数据集GDP-29K
  • 提出结合监督学习和强化学习的训练范式

方法论

结合监督微调与基于可验证奖励的强化学习,训练模型解析几何图形,并用解析结果提升下游任务性能。

原文摘要

Multimodal Large Language Models (MLLMs) have achieved remarkable progress but continue to struggle with geometric reasoning, primarily due to the perception bottleneck regarding fine-grained visual elements. While formal languages have aided plane geometry understanding, solid geometry which requires spatial understanding remains largely unexplored. In this paper, we address this challenge by designing a unified formal language that integrates plane and solid geometry, comprehensively covering geometric structures and semantic relations. We construct GDP-29K, a large-scale dataset comprising 20k plane and 9k solid geometry samples collected from diverse real-world sources, each paired with its ground-truth formal description. To ensure syntactic correctness and geometric consistency, we propose a training paradigm that combines Supervised Fine-Tuning with Reinforcement Learning via Verifiable Rewards. Experiments show that our approach achieves state-of-the-art parsing performance. Furthermore, we demonstrate that our parsed formal descriptions serve as a critical cognitive scaffold, significantly boosting MLLMs' capabilities for downstream geometry reasoning tasks. Our data and code are available at Geoparsing.

标签

几何推理 形式化语言 多模态学习 数据集 强化学习

arXiv 分类

cs.CV