LLM Reasoning 相关度: 8/10

Learning Who Disagrees: Demographic Importance Weighting for Modeling Annotator Distributions with DiADEM

Samay U. Shetty, Tharindu Cyril Weerasooriya, Deepak Pandita, Christopher M. Homan
arXiv: 2604.08425v1 发布: 2026-04-09 更新: 2026-04-09

AI 摘要

DiADEM模型通过学习人口统计学重要性权重,提升了对标注者差异性预测的准确度。

主要贡献

  • 提出了DiADEM神经网络架构
  • 设计了新的item-level差异损失函数
  • 揭示了种族和年龄是驱动标注者差异的关键因素

方法论

DiADEM使用人口统计学投影编码标注者,通过拼接和Hadamard交互融合表示,并使用差异损失进行训练。

原文摘要

When humans label subjective content, they disagree, and that disagreement is not noise. It reflects genuine differences in perspective shaped by annotators' social identities and lived experiences. Yet standard practice still flattens these judgments into a single majority label, and recent LLM-based approaches fare no better: we show that prompted large language models, even with chain-of-thought reasoning, fail to recover the structure of human disagreement. We introduce DiADEM, a neural architecture that learns "how much each demographic axis matters" for predicting who will disagree and on what. DiADEM encodes annotators through per-demographic projections governed by a learned importance vector $\boldsymbolα$, fuses annotator and item representations via complementary concatenation and Hadamard interactions, and is trained with a novel item-level disagreement loss that directly penalizes mispredicted annotation variance. On the DICES conversational-safety and VOICED political-offense benchmarks, DiADEM substantially outperforms both the LLM-as-a-judge and neural model baselines across standard and perspectivist metrics, achieving strong disagreement tracking ($r{=}0.75$ on DICES). The learned $\boldsymbolα$ weights reveal that race and age consistently emerge as the most influential demographic factors driving annotator disagreement across both datasets. Our results demonstrate that explicitly modeling who annotators are not just what they label is essential for NLP systems that aim to faithfully represent human interpretive diversity.

标签

自然语言处理 主观内容理解 人群偏见 模型可解释性

arXiv 分类

cs.AI cs.CL