LLM Memory & RAG 相关度: 9/10

ImplicitMemBench: Measuring Unconscious Behavioral Adaptation in Large Language Models

Chonghan Qin, Xiachong Feng, Weitao Ma, Xiaocheng Feng, Lingpeng Kong
arXiv: 2604.08064v1 发布: 2026-04-09 更新: 2026-04-09

AI 摘要

提出了ImplicitMemBench,用于评估LLM在无意识行为适应方面的能力。

主要贡献

  • 提出了ImplicitMemBench基准,用于评估LLM的内隐记忆
  • 基于认知科学理论,设计了Procedural Memory, Priming, Classical Conditioning三个测试
  • 评估了17个模型,发现LLM在内隐记忆方面存在严重不足

方法论

构建了包含300个条目的测试集,采用Learning/Priming-Interfere-Test协议,以首次尝试的正确率作为评估指标。

原文摘要

Existing memory benchmarks for LLM agents evaluate explicit recall of facts, yet overlook implicit memory where experience becomes automated behavior without conscious retrieval. This gap is critical: effective assistants must automatically apply learned procedures or avoid failed actions without explicit reminders. We introduce ImplicitMemBench, the first systematic benchmark evaluating implicit memory through three cognitively grounded constructs drawn from standard cognitive-science accounts of non-declarative memory: Procedural Memory (one-shot skill acquisition after interference), Priming (theme-driven bias via paired experimental/control instances), and Classical Conditioning (Conditioned Stimulus--Unconditioned Stimulus (CS--US) associations shaping first decisions). Our 300-item suite employs a unified Learning/Priming-Interfere-Test protocol with first-attempt scoring. Evaluation of 17 models reveals severe limitations: no model exceeds 66% overall, with top performers DeepSeek-R1 (65.3%), Qwen3-32B (64.1%), and GPT-5 (63.0%) far below human baselines. Analysis uncovers dramatic asymmetries (inhibition 17.6% vs. preference 75.0%) and universal bottlenecks requiring architectural innovations beyond parameter scaling. ImplicitMemBench reframes evaluation from "what agents recall" to "what they automatically enact".

标签

LLM Implicit Memory Benchmark Cognitive Science

arXiv 分类

cs.AI