Multimodal Learning 相关度: 9/10

CArtBench: Evaluating Vision-Language Models on Chinese Art Understanding, Interpretation, and Authenticity

Xuefeng Wei, Zhixuan Wang, Xuan Zhou, Zhi Qu, Hongyao Li, Yusuke Sakai, Hidetaka Kamigaito, Taro Watanabe
arXiv: 2604.11632v1 发布: 2026-04-13 更新: 2026-04-13

AI 摘要

CArtBench是一个评估VLMs在理解中国艺术品方面的基准测试。

主要贡献

  • 提出了CArtBench基准测试
  • 包含四个子任务:CURATORQA, CATALOGCAPTION, REINTERPRET, CONNOISSEURPAIRS
  • 评估了九个代表性的VLMs在CArtBench上的表现

方法论

通过将维基数据的图像对象与权威目录页面对齐构建数据集,包含五个艺术类别。

原文摘要

We introduce CARTBENCH, a museum-grounded benchmark for evaluating vision-language models (VLMs) on Chinese artworks beyond short-form recognition and QA. CARTBENCH comprises four subtasks: CURATORQA for evidence-grounded recognition and reasoning, CATALOGCAPTION for structured four-section expert-style appreciation, REINTERPRET for defensible reinterpretation with expert ratings, and CONNOISSEURPAIRS for diagnostic authenticity discrimination under visually similar confounds. CARTBENCH is built by aligning image-bearing Palace Museum objects from Wikidata with authoritative catalog pages, spanning five art categories across multiple dynasties. Across nine representative VLMs, we find that high overall CURATORQA accuracy can mask sharp drops on hard evidence linking and style-to-period inference; long-form appreciation remains far from expert references; and authenticity-oriented diagnostic discrimination stays near chance, underscoring the difficulty of connoisseur-level reasoning for current models.

标签

Vision-Language Models Chinese Art Benchmark Evaluation

arXiv 分类

cs.CL