Appear2Meaning: A Cross-Cultural Benchmark for Structured Cultural Metadata Inference from Images
AI 摘要
提出Appear2Meaning跨文化基准,评估VLM在图像中推理结构化文化元数据的能力。
主要贡献
- 提出了跨文化结构化文化元数据推理基准
- 使用LLM-as-Judge框架评估VLM的语义对齐
- 揭示了VLM在跨文化和元数据类型上的性能差异
方法论
构建多类别、跨文化基准数据集,利用LLM作为裁判,评估VLM在精确匹配、部分匹配和属性层面的准确性。
原文摘要
Recent advances in vision-language models (VLMs) have improved image captioning for cultural heritage. However, inferring structured cultural metadata (e.g., creator, origin, period) from visual input remains underexplored. We introduce a multi-category, cross-cultural benchmark for this task and evaluate VLMs using an LLM-as-Judge framework that measures semantic alignment with reference annotations. To assess cultural reasoning, we report exact-match, partial-match, and attribute-level accuracy across cultural regions. Results show that models capture fragmented signals and exhibit substantial performance variation across cultures and metadata types, leading to inconsistent and weakly grounded predictions. These findings highlight the limitations of current VLMs in structured cultural metadata inference beyond visual perception.