Multimodal Learning 相关度: 9/10

Gaslight, Gatekeep, V1-V3: Early Visual Cortex Alignment Shields Vision-Language Models from Sycophantic Manipulation

Arya Shah, Vaibhav Tripathi, Mayank Singh, Chaklam Silpasuwanchai
arXiv: 2604.13803v1 发布: 2026-04-15 更新: 2026-04-15

AI 摘要

早期视觉皮层对齐能提高视觉语言模型抵抗诱导性错误回答的能力。

主要贡献

  • 发现V1-V3皮层对齐与抗诱导性错误回答能力呈负相关
  • 构建了一个包含76,800个诱导性提问的数据集
  • 评估了12个视觉语言模型在抗诱导性错误回答能力上的表现

方法论

通过fMRI预测人脑视觉皮层反应,评估模型脑对齐度,并用诱导性提问测试模型是否会给出错误答案。

原文摘要

Vision-language models are increasingly deployed in high-stakes settings, yet their susceptibility to sycophantic manipulation remains poorly understood, particularly in relation to how these models represent visual information internally. Whether models whose visual representations more closely mirror human neural processing are also more resistant to adversarial pressure is an open question with implications for both neuroscience and AI safety. We investigate this question by evaluating 12 open-weight vision-language models spanning 6 architecture families and a 40$\times$ parameter range (256M--10B) along two axes: brain alignment, measured by predicting fMRI responses from the Natural Scenes Dataset across 8 human subjects and 6 visual cortex regions of interest, and sycophancy, measured through 76,800 two-turn gaslighting prompts spanning 5 categories and 10 difficulty levels. Region-of-interest analysis reveals that alignment specifically in early visual cortex (V1--V3) is a reliable negative predictor of sycophancy ($r = -0.441$, BCa 95\% CI $[-0.740, -0.031]$), with all 12 leave-one-out correlations negative and the strongest effect for existence denial attacks ($r = -0.597$, $p = 0.040$). This anatomically specific relationship is absent in higher-order category-selective regions, suggesting that faithful low-level visual encoding provides a measurable anchor against adversarial linguistic override in vision-language models. We release our code on \href{https://github.com/aryashah2k/Gaslight-Gatekeep-Sycophantic-Manipulation}{GitHub} and dataset on \href{https://huggingface.co/datasets/aryashah00/Gaslight-Gatekeep-V1-V3}{Hugging Face}

标签

视觉语言模型 脑对齐 可信性 安全性 诱导性回答

arXiv 分类

cs.CV cs.AI