Multimodal Learning 相关度: 9/10

Vision-Language Foundation Models for Comprehensive Automated Pavement Condition Assessment

Blessing Agyei Kyem, Joshua Kofi Asamoah, Anthony Dontoh, Armstrong Aboah
arXiv: 2604.08212v1 发布: 2026-04-09 更新: 2026-04-09

AI 摘要

提出PaveGPT,通过领域特定指令调优,实现全面的自动化路面状况评估。

主要贡献

  • 构建PaveInstruct数据集,包含278,889个图像-指令-响应对
  • 训练PaveGPT,在路面状况评估任务上超越现有VLM
  • 验证了指令调优在专业领域VLM中的有效性

方法论

通过统一九个路面数据集的标注,构建PaveInstruct数据集,并在此基础上进行指令调优训练PaveGPT。

原文摘要

General-purpose vision-language models demonstrate strong performance in everyday domains but struggle with specialized technical fields requiring precise terminology, structured reasoning, and adherence to engineering standards. This work addresses whether domain-specific instruction tuning can enable comprehensive pavement condition assessment through vision-language models. PaveInstruct, a dataset containing 278,889 image-instruction-response pairs spanning 32 task types, was created by unifying annotations from nine heterogeneous pavement datasets. PaveGPT, a pavement foundation model trained on this dataset, was evaluated against state-of-the-art vision-language models across perception, understanding, and reasoning tasks. Instruction tuning transformed model capabilities, achieving improvements exceeding 20% in spatial grounding, reasoning, and generation tasks while producing ASTM D6433-compliant outputs. These results enable transportation agencies to deploy unified conversational assessment tools that replace multiple specialized systems, simplifying workflows and reducing technical expertise requirements. The approach establishes a pathway for developing instruction-driven AI systems across infrastructure domains including bridge inspection, railway maintenance, and building condition assessment.

标签

vision-language model pavement condition assessment instruction tuning

arXiv 分类

cs.CV