Entropy-Gradient Grounding: Training-Free Evidence Retrieval in Vision-Language Models
AI 摘要
提出一种无需训练的视觉语言模型证据检索方法,利用熵梯度定位关键视觉区域。
主要贡献
- 提出基于熵梯度的无训练证据检索方法
- 引入迭代缩放和重新对齐过程
- 在多个VLM架构和数据集上验证有效性
方法论
计算模型输出的熵,反向传播到视觉token嵌入,得到熵梯度相关性图,提取相关区域。
原文摘要
Despite rapid progress, pretrained vision-language models still struggle when answers depend on tiny visual details or on combining clues spread across multiple regions, as in documents and compositional queries. We address this by framing grounding as test-time evidence retrieval: given a query, the model should actively identify where to look next to resolve ambiguity. To this end, we propose a training-free, model-intrinsic grounding method that uses uncertainty as supervision. Specifically, we compute the entropy of the model's next-token distribution and backpropagate it to the visual token embeddings to obtain an entropy-gradient relevance map, without auxiliary detectors or attention-map heuristics. We then extract and rank multiple coherent regions to support multi-evidence queries, and introduce an iterative zoom-and-reground procedure with a spatial-entropy stopping rule to avoid over-refinement. Experiments on seven benchmarks across four VLM architectures demonstrate consistent improvements over existing methods, with the largest gains on detail-critical and high-resolution settings, while also producing more interpretable evidence localizations.