LLM Memory & RAG 相关度: 9/10

A-MBER: Affective Memory Benchmark for Emotion Recognition

Deliang Wen, Ke Sun, Yu Wang
arXiv: 2604.07017v1 发布: 2026-04-08 更新: 2026-04-08

AI 摘要

A-MBER是一个情感记忆基准,用于评估模型利用交互历史理解用户情感状态的能力。

主要贡献

  • 提出了A-MBER情感记忆基准数据集。
  • A-MBER数据集侧重于基于多轮交互历史的情感推断。
  • 设计了判断、检索和解释等任务来评估模型能力。

方法论

通过多阶段流程构建A-MBER数据集,包括规划、对话生成、标注、问题构建和最终打包,并支持多种评估任务和鲁棒性测试。

原文摘要

AI assistants that interact with users over time need to interpret the user's current emotional state in order to respond appropriately and personally. However, this capability remains insufficiently evaluated. Existing emotion datasets mainly assess local or instantaneous affect, while long-term memory benchmarks focus largely on factual recall, temporal consistency, or knowledge updating. As a result, current resources provide limited support for testing whether a model can use remembered interaction history to interpret a user's present affective state. We introduce A-MBER, an Affective Memory Benchmark for Emotion Recognition, to evaluate this capability. A-MBER focuses on present affective interpretation grounded in remembered multi-session interaction history. Given an interaction trajectory and a designated anchor turn, a model must infer the user's current affective state, identify historically relevant evidence, and justify its interpretation in a grounded way. The benchmark is constructed through a staged pipeline with explicit intermediate representations, including long-horizon planning, conversation generation, annotation, question construction, and final packaging. It supports judgment, retrieval, and explanation tasks, together with robustness settings such as modality degradation and insufficient-evidence conditions. Experiments compare local-context, long-context, retrieved-memory, structured-memory, and gold-evidence conditions within a unified framework. Results show that A-MBER is especially discriminative on the subsets it is designed to stress, including long-range implicit affect, high-dependency memory levels, trajectory-based reasoning, and adversarial settings. These findings suggest that memory supports affective interpretation not simply by providing more history, but by enabling more selective, grounded, and context-sensitive use of past interaction

标签

emotion recognition long-term memory benchmark affective computing interaction history

arXiv 分类

cs.AI