Multimodal Learning 相关度: 9/10

Learning to Think Like a Cartoon Captionist: Incongruity-Resolution Supervision for Multimodal Humor Understanding

Hatice Merve Vural, Doga Kukul, Ege Erdem Ozlu, Demir Ekin Arikan, Bob Mankoff, Erkut Erdem, Aykut Erdem
arXiv: 2604.15210v1 发布: 2026-04-16 更新: 2026-04-16

AI 摘要

该论文提出IRS框架,通过监督中间推理过程,提升模型在多模态幽默理解任务上的性能,超越现有模型。

主要贡献

  • 提出IRS框架,分解幽默理解为不协调建模、解决建模和偏好对齐三个组件
  • 通过结构化traces监督中间推理过程,使模型学习从视觉感知到幽默解释的路径
  • 在NYCC数据集上,IRS框架显著优于现有基线模型,并在外部基准测试中表现出良好的泛化能力

方法论

IRS框架基于不协调解决理论和专家标题撰写实践,通过监督中间推理过程来指导模型学习幽默理解。

原文摘要

Humor is one of the few cognitive tasks where getting the reasoning right matters as much as getting the answer right. While recent work evaluates humor understanding on benchmarks such as the New Yorker Cartoon Caption Contest (NYCC), it largely treats it as black-box prediction, overlooking the structured reasoning processes underlying humor comprehension. We introduce IRS (Incongruity-Resolution Supervision), a framework that decomposes humor understanding into three components: incongruity modeling, which identifies mismatches in the visual scene; resolution modeling, which constructs coherent reinterpretations of these mismatches; and preference alignment, which evaluates candidate interpretations under human judgments. Grounded in incongruity-resolution theory and expert captionist practice, IRS supervises intermediate reasoning process through structured traces that make the path from visual perception to humorous interpretation explicit and learnable. Across 7B, 32B, and 72B models on NYCC, IRS outperforms strong open and closed multimodal baselines across caption matching and ranking tasks, with our largest model approaching expert-level performance on ranking. Zero-shot transfer to external benchmarks shows that IRS learns generalizable reasoning patterns. Our results suggest that supervising reasoning structure, rather than scale alone, is key for reasoning-centric tasks.

标签

multimodal humor understanding incongruity-resolution reasoning supervision

arXiv 分类

cs.AI cs.CL