Continuous Interpretive Steering for Scalar Diversity
AI 摘要
该论文提出了一种评估LLM中细粒度语用推理能力的方法CIS,并构建了数据集GraSD。
主要贡献
- 提出了Continuous Interpretive Steering (CIS) 方法,用于探测LLM中的语用推理能力。
- 构建了GraSD数据集,用于评估LLM在标量多样性上的表现。
- 实验结果表明,分级的激活控制可以恢复LLM中编码的语用敏感性。
方法论
通过连续改变激活层的steering强度,观察LLM对不同标量项的语用推理变化,并使用GraSD数据集进行评估。
原文摘要
Pragmatic inference is inherently graded. Different lexical items give rise to pragmatic enrichment to different degrees. Scalar implicature exemplifies this property through scalar diversity, where implicature strength varies across scalar items. However, evaluations of pragmatic inference in large language models (LLMs) often rely on prompt-based manipulations. Beyond prompt-level effects, this study introduces Continuous Interpretive Steering (CIS), a method that probes graded pragmatic interpretation by treating activation-level steering strength as a continuous experimental variable. To support this analysis, this study introduces a new dataset, GraSD, which encodes graded scalar diversity. Experiments on four LLMs show that uniform activation steering increases pragmatic interpretations globally but collapses item-level variation, whereas graded activation steering yields differentiated interpretive shifts aligned with scalar diversity grades. It indicates that graded sensitivity is encoded in the representation space and can be systematically recovered through controlled intervention. Together, CIS and GraSD provide a principled framework for evaluating graded pragmatic sensitivity in LLMs.