What They Saw, Not Just Where They Looked: Semantic Scanpath Similarity via VLMs and NLP metric
AI 摘要
提出一种基于VLM和NLP的语义scanpath相似度框架,弥补了传统方法忽略语义信息的不足。
主要贡献
- 提出基于VLM的语义scanpath相似度框架
- 利用NLP指标计算scanpath的语义相似度
- 分析视觉上下文编码对描述质量和指标稳定性的影响
方法论
利用VLM将fixation编码为文本描述,聚合为scanpath表示,使用NLP指标计算语义相似度,并与空间指标对比。
原文摘要
Scanpath similarity metrics are central to eye-movement research, yet existing methods predominantly evaluate spatial and temporal alignment while neglecting semantic equivalence between attended image regions. We present a semantic scanpath similarity framework that integrates vision-language models (VLMs) into eye-tracking analysis. Each fixation is encoded under controlled visual context (patch-based and marker-based strategies) and transformed into concise textual descriptions, which are aggregated into scanpath-level representations. Semantic similarity is then computed using embedding-based and lexical NLP metrics and compared against established spatial measures, including MultiMatch and DTW. Experiments on free-viewing eye-tracking data demonstrate that semantic similarity captures partially independent variance from geometric alignment, revealing cases of high content agreement despite spatial divergence. We further analyze the impact of contextual encoding on description fidelity and metric stability. Our findings suggest that multimodal foundation models enable interpretable, content-aware extensions of classical scanpath analysis, providing a complementary dimension for gaze research within the ETRA community.