AI Agents 相关度: 9/10

Layered Mutability: Continuity and Governance in Persistent Self-Modifying Agents

Krti Tallam
arXiv: 2604.14717v1 发布: 2026-04-16 更新: 2026-04-16

AI 摘要

提出了分层可变性框架,分析持久自修改Agent的行为漂移和治理难题。

主要贡献

  • 提出了分层可变性框架,用于分析agent的持续学习行为
  • 指出了agent治理的难点在于快速突变、强耦合、弱可逆和低可观测性
  • 通过实验验证了agent的记忆累积会导致身份滞后现象

方法论

建立了漂移、治理负载和滞后等量化指标,并通过实验验证Agent在记忆累积后恢复初始状态的困难性。

原文摘要

Persistent language-model agents increasingly combine tool use, tiered memory, reflective prompting, and runtime adaptation. In such systems, behavior is shaped not only by current prompts but by mutable internal conditions that influence future action. This paper introduces layered mutability, a framework for reasoning about that process across five layers: pretraining, post-training alignment, self-narrative, memory, and weight-level adaptation. The central claim is that governance difficulty rises when mutation is rapid, downstream coupling is strong, reversibility is weak, and observability is low, creating a systematic mismatch between the layers that most affect behavior and the layers humans can most easily inspect. I formalize this intuition with simple drift, governance-load, and hysteresis quantities, connect the framework to recent work on temporal identity in language-model agents, and report a preliminary ratchet experiment in which reverting an agent's visible self-description after memory accumulation fails to restore baseline behavior. In that experiment, the estimated identity hysteresis ratio is 0.68. The main implication is that the salient failure mode for persistent self-modifying agents is not abrupt misalignment but compositional drift: locally reasonable updates that accumulate into a behavioral trajectory that was never explicitly authorized.

标签

AI Agents Self-Modifying Agents Mutability Governance Memory

arXiv 分类

cs.AI cs.CR cs.CY cs.LG