Multimodal Learning 相关度: 9/10

AICA-Bench: Holistically Examining the Capabilities of VLMs in Affective Image Content Analysis

Dong She, Xianrong Yao, Liqun Chen, Jinghe Yu, Yang Gao, Zhanpeng Jin
arXiv: 2604.05900v1 发布: 2026-04-07 更新: 2026-04-07

AI 摘要

AICA-Bench基准测试VLMs在情感图像内容分析的能力,并提出GAT Prompting框架。

主要贡献

  • 提出了AICA-Bench基准测试
  • 分析了VLMs在情感图像内容分析中的局限性
  • 提出了Grounded Affective Tree (GAT) Prompting框架

方法论

构建包含三个任务的基准测试评估VLMs,并设计结合视觉支架和分层推理的Prompting框架。

原文摘要

Vision-Language Models (VLMs) have demonstrated strong capabilities in perception, yet holistic Affective Image Content Analysis (AICA), which integrates perception, reasoning, and generation into a unified framework, remains underexplored. To address this gap, we introduce AICA-Bench, a comprehensive benchmark with three core tasks: Emotion Understanding (EU), Emotion Reasoning (ER), and Emotion-Guided Content Generation (EGCG). We evaluate 23 VLMs and identify two major limitations: weak intensity calibration and shallow open-ended descriptions. To address these issues, we propose Grounded Affective Tree (GAT) Prompting, a training-free framework that combines visual scaffolding with hierarchical reasoning. Experiments show that GAT reduces intensity errors and improves descriptive depth, providing a strong baseline for future research on affective multimodal understanding and generation.

标签

VLMs 情感分析 基准测试 Prompting

arXiv 分类

cs.CV