Multimodal Learning 相关度: 9/10

Vision-Guided Iterative Refinement for Frontend Code Generation

Hannah Sansford, Derek H. C. Law, Wei Liu, Abhishek Tripathi, Niresh Agarwal, Gerrit J. J. van den Burg
arXiv: 2604.05839v1 发布: 2026-04-07 更新: 2026-04-07

AI 摘要

提出一种基于视觉语言模型的全自动前端代码生成迭代优化框架,显著提高代码质量。

主要贡献

  • 提出基于视觉语言模型的自动代码评审和迭代优化框架
  • 验证该框架在前端代码生成任务上的有效性,显著提升代码质量
  • 研究参数高效微调方法,将视觉评审能力内化到代码生成模型中

方法论

利用视觉语言模型作为视觉评审器,提供结构化反馈,指导代码生成模型的迭代优化,并使用LoRA进行参数高效微调。

原文摘要

Code generation with large language models often relies on multi-stage human-in-the-loop refinement, which is effective but very costly - particularly in domains such as frontend web development where the solution quality depends on rendered visual output. We present a fully automated critic-in-the-loop framework in which a vision-language model serves as a visual critic that provides structured feedback on rendered webpages to guide iterative refinement of generated code. Across real-world user requests from the WebDev Arena dataset, this approach yields consistent improvements in solution quality, achieving up to 17.8% increase in performance over three refinement cycles. Next, we investigate parameter-efficient fine-tuning using LoRA to understand whether the improvements provided by the critic can be internalized by the code-generating LLM. Fine-tuning achieves 25% of the gains from the best critic-in-the-loop solution without a significant increase in token counts. Our findings indicate that automated, VLM-based critique of frontend code generation leads to significantly higher quality solutions than can be achieved through a single LLM inference pass, and highlight the importance of iterative refinement for the complex visual outputs associated with web development.

标签

代码生成 视觉语言模型 迭代优化 前端开发

arXiv 分类

cs.AI