Multimodal Learning 相关度: 9/10

"I See What You Did There": Can Large Vision-Language Models Understand Multimodal Puns?

Naen Xu, Jiayi Sheng, Changjiang Li, Chunyi Zhou, Yuyuan Li, Tianyu Du, Jun Wang, Zhihui Fu, Jinbao Li, Shouling Ji
arXiv: 2604.05930v1 发布: 2026-04-07 更新: 2026-04-07

AI 摘要

该论文研究了视觉-语言模型理解多模态双关语的能力,并提出了数据集和改进方法。

主要贡献

  • 提出了多模态双关语生成流程
  • 构建了包含多种双关语和对抗性非双关语的数据集MultiPun
  • 提出了增强双关语理解的prompt和模型层面策略

方法论

构建双关语生成流程和数据集,评估现有VLM,并提出prompt和模型层面改进策略,使用F1分数评估。

原文摘要

Puns are a common form of rhetorical wordplay that exploits polysemy and phonetic similarity to create humor. In multimodal puns, visual and textual elements synergize to ground the literal sense and evoke the figurative meaning simultaneously. Although Vision-Language Models (VLMs) are widely used in multimodal understanding and generation, their ability to understand puns has not been systematically studied due to a scarcity of rigorous benchmarks. To address this, we first propose a multimodal pun generation pipeline. We then introduce MultiPun, a dataset comprising diverse types of puns alongside adversarial non-pun distractors. Our evaluation reveals that most models struggle to distinguish genuine puns from these distractors. Moreover, we propose both prompt-level and model-level strategies to enhance pun comprehension, with an average improvement of 16.5% in F1 scores. Our findings provide valuable insights for developing future VLMs that master the subtleties of human-like humor via cross-modal reasoning.

标签

多模态 双关语 视觉-语言模型 数据集 幽默

arXiv 分类

cs.CL cs.AI