Multimodal Learning 相关度: 8/10

Learn to Rank: Visual Attribution by Learning Importance Ranking

David Schinagl, Christian Fruhwirth-Reisinger, Alexander Prutsch, Samuel Schulter, Horst Possegger
arXiv: 2604.05819v1 发布: 2026-04-07 更新: 2026-04-07

AI 摘要

提出了一种直接优化deletion和insertion指标的学习排序视觉归因方法,提升了Transformer模型的可解释性。

主要贡献

  • 提出直接优化deletion/insertion指标的学习框架
  • 使用Gumbel-Sinkhorn松弛实现可微排序
  • 单次前向传递生成像素级归因,可选梯度优化

方法论

将视觉归因问题转化为排序学习问题,利用Gumbel-Sinkhorn松弛实现可微分排序,直接优化deletion和insertion指标。

原文摘要

Interpreting the decisions of complex computer vision models is crucial to establish trust and accountability, especially in safety-critical domains. An established approach to interpretability is generating visual attribution maps that highlight regions of the input most relevant to the model's prediction. However, existing methods face a three-way trade-off. Propagation-based approaches are efficient, but they can be biased and architecture-specific. Meanwhile, perturbation-based methods are causally grounded, yet they are expensive and for vision transformers often yield coarse, patch-level explanations. Learning-based explainers are fast but usually optimize surrogate objectives or distill from heuristic teachers. We propose a learning scheme that instead optimizes deletion and insertion metrics directly. Since these metrics depend on non-differentiable sorting and ranking, we frame them as permutation learning and replace the hard sorting with a differentiable relaxation using Gumbel-Sinkhorn. This enables end-to-end training through attribution-guided perturbations of the target model. During inference, our method produces dense, pixel-level attributions in a single forward pass with optional, few-step gradient refinement. Our experiments demonstrate consistent quantitative improvements and sharper, boundary-aligned explanations, particularly for transformer-based vision models.

标签

视觉归因 可解释性 Transformer 排序学习

arXiv 分类

cs.CV cs.LG