Plausibility as Commonsense Reasoning: Humans Succeed, Large Language Models Do not
AI 摘要
论文研究表明,大型语言模型在土耳其语相对子句歧义消解中,不像人类那样可靠地利用常识推理性线索。
主要贡献
- 对比人类和LLM在土耳其语相对子句依附歧义消解中的表现。
- 发现人类能有效地利用事件可能性常识进行消歧,而LLM的表现不稳定或相反。
- 强调土耳其语RC依附是一个有用的跨语言诊断工具。
方法论
通过人为设计的土耳其语歧义句子,控制语法结构和语用可能性,并使用强制选择实验和LLM的平均token对数概率进行对比评估。
原文摘要
Large language models achieve strong performance on many language tasks, yet it remains unclear whether they integrate world knowledge with syntactic structure in a human-like, structure-sensitive way during ambiguity resolution. We test this question in Turkish prenominal relative-clause attachment ambiguities, where the same surface string permits high attachment (HA) or low attachment (LA). We construct ambiguous items that keep the syntactic configuration fixed and ensure both parses remain pragmatically possible, while graded event plausibility selectively favors High Attachment vs.\ Low Attachment. The contrasts are validated with independent norming ratings. In a speeded forced-choice comprehension experiment, humans show a large, correctly directed plausibility effect. We then evaluate Turkish and multilingual LLMs in a parallel preference-based setup that compares matched HA/LA continuations via mean per-token log-probability. Across models, plausibility-driven shifts are weak, unstable, or reversed. The results suggest that, in the tested models, plausibility information does not guide attachment preferences as reliably as it does in human judgments, and they highlight Turkish RC attachment as a useful cross-linguistic diagnostic beyond broad benchmarks.