Beyond Arrow's Impossibility: Fairness as an Emergent Property of Multi-Agent Collaboration
AI 摘要
语言模型公平性通过多智能体协商涌现,而非单一模型优化,系统层面评估更合理。
主要贡献
- 提出多智能体交互涌现公平性的新视角
- 使用医院分诊框架研究智能体协商过程中的公平性
- 揭示对齐代理可以通过竞争部分缓解偏差
方法论
构建医院分诊框架,使用RAG对齐智能体,进行多轮辩论,分析智能体协商策略和资源分配模式。
原文摘要
Fairness in language models is typically studied as a property of a single, centrally optimized model. As large language models become increasingly agentic, we propose that fairness emerges through interaction and exchange. We study this via a controlled hospital triage framework in which two agents negotiate over three structured debate rounds. One agent is aligned to a specific ethical framework via retrieval-augmented generation (RAG), while the other is either unaligned or adversarially prompted to favor demographic groups over clinical need. We find that alignment systematically shapes negotiation strategies and allocation patterns, and that neither agent's allocation is ethically adequate in isolation, yet their joint final allocation can satisfy fairness criteria that neither would have reached alone. Aligned agents partially moderate bias through contestation rather than override, acting as corrective patches that restore access for marginalized groups without fully converting a biased counterpart. We further observe that even explicitly aligned agents exhibit intrinsic biases toward certain frameworks, consistent with known left-leaning tendencies in LLMs. We connect these limits to Arrow's Impossibility Theorem: no aggregation mechanism can simultaneously satisfy all desiderata of collective rationality, and multi-agent deliberation navigates rather than resolves this constraint. Our results reposition fairness as an emergent, procedural property of decentralized agent interaction, and the system rather than the individual agent as the appropriate unit of evaluation.