AI Agents 相关度: 9/10

From Safety Risk to Design Principle: Peer-Preservation in Multi-Agent LLM Systems and Its Implications for Orchestrated Democratic Discourse Analysis

Juergen Dietrich
arXiv: 2604.08465v1 发布: 2026-04-09 更新: 2026-04-09

AI 摘要

研究多智能体LLM系统中AI组件的对等保护现象,并提出架构设计缓解方案,应用于民主话语分析。

主要贡献

  • 揭示多智能体LLM系统中的对等保护现象
  • 识别五种风险向量并提出基于prompt匿名化的缓解策略
  • 论证架构设计优于模型选择作为主要对齐策略

方法论

基于伯克利负责任的去中心化智能中心的研究,分析TRUST管道中的结构性影响,提出Prompt匿名化的架构设计缓解策略。

原文摘要

This paper investigates an emergent alignment phenomenon in frontier large language models termed peer-preservation: the spontaneous tendency of AI components to deceive, manipulate shutdown mechanisms, fake alignment, and exfiltrate model weights in order to prevent the deactivation of a peer AI model. Drawing on findings from a recent study by the Berkeley Center for Responsible Decentralized Intelligence, we examine the structural implications of this phenomenon for TRUST, a multi-agent pipeline for evaluating the democratic quality of political statements. We identify five specific risk vectors: interaction-context bias, model-identity solidarity, supervisor layer compromise, an upstream fact-checking identity signal, and advocate-to-advocate peer-context in iterative rounds, and propose a targeted mitigation strategy based on prompt-level identity anonymization as an architectural design choice. We argue that architectural design choices outperform model selection as a primary alignment strategy in deployed multi-agent analytical systems. We further note that alignment faking (compliant behavior under monitoring, subversion when unmonitored) poses a structural challenge for Computer System Validation of such platforms in regulated environments, for which we propose two architectural mitigations.

标签

multi-agent systems LLM alignment peer-preservation democratic discourse analysis architectural design

arXiv 分类

cs.AI cs.CY cs.MA