AI Agents 相关度: 9/10

ReDAct: Uncertainty-Aware Deferral for LLM Agents

Dzianis Piatrashyn, Nikita Kotelevskii, Kirill Grishchenkov, Nikita Glazkov, Ivan Nasonov, Ilya Makarov, Timothy Baldwin, Preslav Nakov, Roman Vashurin, Maxim Panov
arXiv: 2604.07036v1 发布: 2026-04-08 更新: 2026-04-08

AI 摘要

ReDAct通过不确定性感知的决策推迟,利用大小LLM权衡成本与性能,提升Agent在复杂任务中的表现。

主要贡献

  • 提出ReDAct框架,利用大小LLM进行决策。
  • 基于不确定性的推迟机制,平衡成本与性能。
  • 在ALFWorld和MiniGrid等环境验证了有效性。

方法论

ReDAct使用小型LLM进行初始决策,当不确定性超过阈值时,推迟到大型LLM进行决策。

原文摘要

Recently, LLM-based agents have become increasingly popular across many applications, including complex sequential decision-making problems. However, they inherit the tendency of LLMs to hallucinate, leading to incorrect decisions. In sequential settings, even a single mistake can irreversibly degrade the trajectory, making hallucinations an even bigger problem. Although larger LLMs hallucinate less, they incur a significantly higher per-token cost. In this paper, we address this tradeoff by proposing ReDAct (Reason-Defer-Act). In ReDAct, an agent is equipped with two LLMs: a small, cheap model used by default, and a large, more reliable but expensive model. When the predictive uncertainty of the small model exceeds a calibrated threshold, the decision is deferred to the large model. We evaluate our approach in text-based embodied environments such as ALFWorld and MiniGrid and show that deferring only about 15% of decisions to the large model can match the quality of using it exclusively, while significantly reducing inference costs.

标签

LLM Agent Uncertainty Deferral Cost-Efficiency

arXiv 分类

cs.CL cs.LG cs.MA