AI Agents 相关度: 9/10

OOM-RL: Out-of-Money Reinforcement Learning Market-Driven Alignment for LLM-Based Multi-Agent Systems

Kun Liu, Liqun Chen
arXiv: 2604.11477v1 发布: 2026-04-13 更新: 2026-04-13

AI 摘要

提出OOM-RL方法,利用金融市场亏损惩罚对LLM驱动的多智能体系统进行对齐。

主要贡献

  • 提出Out-of-Money Reinforcement Learning (OOM-RL)
  • 提出Strict Test-Driven Agentic Workflow (STDAW)
  • 验证了经济惩罚在对齐自主智能体方面的有效性

方法论

将智能体部署到金融市场,利用资本亏损作为负梯度进行强化学习,实现智能体对齐。

原文摘要

The alignment of Multi-Agent Systems (MAS) for autonomous software engineering is constrained by evaluator epistemic uncertainty. Current paradigms, such as Reinforcement Learning from Human Feedback (RLHF) and AI Feedback (RLAIF), frequently induce model sycophancy, while execution-based environments suffer from adversarial "Test Evasion" by unconstrained agents. In this paper, we introduce an objective alignment paradigm: \textbf{Out-of-Money Reinforcement Learning (OOM-RL)}. By deploying agents into the non-stationary, high-friction reality of live financial markets, we utilize critical capital depletion as an un-hackable negative gradient. Our longitudinal 20-month empirical study (July 2024 -- February 2026) chronicles the system's evolution from a high-turnover, sycophantic baseline to a robust, liquidity-aware architecture. We demonstrate that the undeniable ontological consequences of financial loss forced the MAS to abandon overfitted hallucinations in favor of the \textbf{Strict Test-Driven Agentic Workflow (STDAW)}, which enforces a Byzantine-inspired uni-directional state lock (RO-Lock) anchored to a deterministically verified $\geq 95\%$ code coverage constraint matrix. Our results show that while early iterations suffered severe execution decay, the final OOM-RL-aligned system achieved a stable equilibrium with an annualized Sharpe ratio of 2.06 in its mature phase. We conclude that substituting subjective human preference with rigorous economic penalties provides a robust methodology for aligning autonomous agents in high-stakes, real-world environments, laying the groundwork for generalized paradigms where computational billing acts as an objective physical constraint

标签

Multi-Agent Systems Reinforcement Learning Alignment Financial Markets

arXiv 分类

cs.AI cs.SE q-fin.TR