Multimodal Learning 相关度: 9/10

Why Do Vision Language Models Struggle To Recognize Human Emotions?

Madhav Agarwal, Sotirios A. Tsaftaris, Laura Sevilla-Lara, Steven McDonagh
arXiv: 2604.15280v1 发布: 2026-04-16 更新: 2026-04-16

AI 摘要

分析了VLM在识别人类情感方面的不足,并提出了针对长尾分布和时间信息缺失的解决方案。

主要贡献

  • 发现VLM在情感识别中存在的长尾分布偏见问题
  • 揭示VLM稀疏时间采样与微表情的短暂性不匹配
  • 提出多阶段上下文丰富策略,利用帧间信息增强情感识别

方法论

通过观察和分析,发现VLM的脆弱性,并提出基于采样策略和上下文信息增强的方法,验证了其有效性。

原文摘要

Understanding emotions is a fundamental ability for intelligent systems to be able to interact with humans. Vision-language models (VLMs) have made tremendous progress in the last few years for many visual tasks, potentially offering a promising solution for understanding emotions. However, it is surprising that even the most sophisticated contemporary VLMs struggle to recognize human emotions or to outperform even specialized vision-only classifiers. In this paper we ask the question "Why do VLMs struggle to recognize human emotions?", and observe that the inherently continuous and dynamic task of facial expression recognition (DFER) exposes two critical VLM vulnerabilities. First, emotion datasets are naturally long-tailed, and the web-scale data used to pre-train VLMs exacerbates this head-class bias, causing them to systematically collapse rare, under-represented emotions into common categories. We propose alternative sampling strategies that prevent favoring common concepts. Second, temporal information is critical for understanding emotions. However, VLMs are unable to represent temporal information over dense frame sequences, as they are limited by context size and the number of tokens that can fit in memory, which poses a clear challenge for emotion recognition. We demonstrate that the sparse temporal sampling strategy used in VLMs is inherently misaligned with the fleeting nature of micro-expressions (0.25-0.5 seconds), which are often the most critical affective signal. As a diagnostic probe, we propose a multi-stage context enrichment strategy that utilizes the information from "in-between" frames by first converting them into natural language summaries. This enriched textual context is provided as input to the VLM alongside sparse keyframes, preventing attentional dilution from excessive visual data while preserving the emotional trajectory.

标签

Vision-Language Models Emotion Recognition Long-Tailed Distribution Temporal Information

arXiv 分类

cs.CV cs.AI