Multimodal Learning 相关度: 7/10

From Attribution to Action: A Human-Centered Application of Activation Steering

Tobias Labarta, Maximilian Dreyer, Katharina Weitz, Wojciech Samek, Sebastian Lapuschkin
arXiv: 2604.11467v1 发布: 2026-04-13 更新: 2026-04-13

AI 摘要

该论文研究通过激活控制使XAI更具操作性,并探讨其在视觉模型调试中的应用。

主要贡献

  • 提出一个交互式工作流,结合SAE归因和激活控制
  • 通过用户研究分析了激活控制在模型调试中的应用效果
  • 揭示了激活控制的优势和风险,为安全有效使用提供参考

方法论

该论文采用混合方法,结合了技术开发(交互式工具),用户研究(专家访谈),以及定性和定量分析。

原文摘要

Explainable AI (XAI) methods reveal which features influence model predictions, yet provide limited means for practitioners to act on these explanations. Activation steering of components identified via XAI offers a path toward actionable explanations, although its practical utility remains understudied. We introduce an interactive workflow combining SAE-based attribution with activation steering for instance-level analysis of concept usage in vision models, implemented as a web-based tool. Based on this workflow, we conduct semi-structured expert interviews (N=8) with debugging tasks on CLIP to investigate how practitioners reason about, trust, and apply activation steering. We find that steering enables a shift from inspection to intervention-based hypothesis testing (8/8 participants), with most grounding trust in observed model responses rather than explanation plausibility alone (6/8). Participants adopted systematic debugging strategies dominated by component suppression (7/8) and highlighted risks including ripple effects and limited generalization of instance-level corrections. Overall, activation steering renders interpretability more actionable while raising important considerations for safe and effective use.

标签

XAI Activation Steering Interpretability Debugging Vision Models

arXiv 分类

cs.AI cs.HC cs.LG