LLM Reasoning 相关度: 9/10

METER: Evaluating Multi-Level Contextual Causal Reasoning in Large Language Models

Pengfeng Li, Chen Huang, Chaoqun Hao, Hongyao Chen, Xiao-Yong Wei, Wenqiang Lei, See-Kiong Ng
arXiv: 2604.11502v1 发布: 2026-04-13 更新: 2026-04-13

AI 摘要

该论文提出了METER基准,用于评估LLM在多层次因果推理中的能力,并分析了其失效模式。

主要贡献

  • 提出了METER基准用于评估LLM的多层次因果推理能力
  • 揭示了LLM在因果推理层次上升时性能显著下降
  • 分析了LLM因果推理的两种主要失效模式:信息干扰和上下文忠实度下降

方法论

通过构建统一上下文设置下的因果推理基准,对不同LLM进行评估,并通过错误模式识别和内部信息流追踪进行机制分析。

原文摘要

Contextual causal reasoning is a critical yet challenging capability for Large Language Models (LLMs). Existing benchmarks, however, often evaluate this skill in fragmented settings, failing to ensure context consistency or cover the full causal hierarchy. To address this, we pioneer METER to systematically benchmark LLMs across all three levels of the causal ladder under a unified context setting. Our extensive evaluation of various LLMs reveals a significant decline in proficiency as tasks ascend the causal hierarchy. To diagnose this degradation, we conduct a deep mechanistic analysis via both error pattern identification and internal information flow tracing. Our analysis reveals two primary failure modes: (1) LLMs are susceptible to distraction by causally irrelevant but factually correct information at lower level of causality; and (2) as tasks ascend the causal hierarchy, faithfulness to the provided context degrades, leading to a reduced performance. We belive our work advances our understanding of the mechanisms behind LLM contextual causal reasoning and establishes a critical foundation for future research. Our code and dataset are available at https://github.com/SCUNLP/METER .

标签

因果推理 大语言模型 评估基准 机制分析

arXiv 分类

cs.CL cs.AI