Multimodal Learning 相关度: 9/10

Scientific Graphics Program Synthesis via Dual Self-Consistency Reinforcement Learning

Juekai Lin, Yun Zhu, Honglin Lin, Sijing Li, Tianwei Lin, Zheng Liu, Xiaoyang Wang, Wenqiao Zhang, Lijun Wu
arXiv: 2604.06079v1 发布: 2026-04-07 更新: 2026-04-07

AI 摘要

该论文提出了一种用于科学图形程序合成的自洽强化学习框架,并构建了高质量数据集和基准。

主要贡献

  • 构建了高质量大型数据集SciTikZ-230K
  • 提出了多方面的基准SciTikZ-Bench
  • 引入了双自洽强化学习优化范式

方法论

采用执行中心的数据引擎构建数据集,并使用双自洽强化学习优化模型,通过往返验证提高代码质量。

原文摘要

Graphics Program Synthesis is pivotal for interpreting and editing visual data, effectively facilitating the reverse-engineering of static visuals into editable TikZ code. While TikZ is the de facto standard for scientific schematics due to its programmatic flexibility, its requirement for rigorous spatial precision presents a significant challenge for Multimodal Large Language Models. Progress is currently stifled by two primary gaps: (1) Data Quality Gap: existing image-TikZ corpora often lack strict executability and reliable visual alignment; (2) Evaluation Gap: a lack of benchmarks for both structural and visual fidelity. To address these, we present a closed-loop framework featuring: SciTikZ-230K, a large-scale, high-quality dataset from our Execution-Centric Data Engine covering 11 diverse scientific disciplines; SciTikZ-Bench, a multifaceted benchmark spanning from basic geometric constructs to intricate hierarchical schematics to evaluate both visual fidelity and structural logic. To further broaden the scope of visual-code optimization methodology, we introduce a novel Dual Self-Consistency Reinforcement Learning optimization paradigm, which utilizes Round-Trip Verification to penalize degenerate code and boost overall self-consistency. Empowered by these, our trained model SciTikZer-8B achieves state-of-the-art performance, consistently outperforming proprietary giants like Gemini-2.5-Pro and massive models like Qwen3-VL-235B-A22B-Instruct.

标签

图形程序合成 强化学习 多模态 科学绘图

arXiv 分类

cs.CV cs.AI