AI Agents 相关度: 9/10

Select-then-Solve: Paradigm Routing as Inference-Time Optimization for LLM Agents

Heng Zhou, Zelin Tan, Zhemeng Zhang, Yutao Fan, Yibing Lin, Li Kang, Xiufeng Song, Rui Li, Songtao Huang, Ao Yu, Yuchen Fan, Yanxu Chen, Kaixin Xu, Xiaohong Liu, Yiran Qin, Philip Torr, Chen Zhang, Zhenfei Yin
arXiv: 2604.06753v1 发布: 2026-04-08 更新: 2026-04-08

AI 摘要

该论文提出一种基于嵌入的路由方法,为LLM Agent选择最优推理范式,显著提升性能。

主要贡献

  • 对比六种推理范式在不同LLM和任务上的表现,发现没有单一范式占优
  • 提出Select-then-Solve方法,使用轻量级路由选择最佳范式
  • 实验证明该方法优于固定范式,并缩小了与oracle选择的差距

方法论

通过嵌入式路由根据任务特征选择最合适的推理范式,并在下游任务上进行评估和比较。

原文摘要

When an LLM-based agent improves on a task, is the gain from the model itself or from the reasoning paradigm wrapped around it? We study this question by comparing six inference-time paradigms, namely Direct, CoT, ReAct, Plan-Execute, Reflection, and ReCode, across four frontier LLMs and ten benchmarks, yielding roughly 18,000 runs. We find that reasoning structure helps dramatically on some tasks but hurts on others: ReAct improves over Direct by 44pp on GAIA, while CoT degrades performance by 15pp on HumanEval. No single paradigm dominates, and oracle per-task selection beats the best fixed paradigm by 17.1pp on average. Motivated by this complementarity, we propose a select-then-solve approach: before answering each task, a lightweight embedding-based router selects the most suitable paradigm. Across four models, the router improves average accuracy from 47.6% to 53.1%, outperforming the best fixed paradigm at 50.3% by 2.8pp and recovering up to 37% of the oracle gap. In contrast, zero-shot self-routing only works for GPT-5 at 67.1% and fails for weaker models, all trailing the learned router. Our results argue that reasoning paradigm selection should be a per-task decision made by a learned router, not a fixed architectural choice.

标签

LLM Agents Reasoning Inference-time Optimization Paradigm Selection

arXiv 分类

cs.CL