LLM Reasoning 相关度: 9/10

When to Trust Tools? Adaptive Tool Trust Calibration For Tool-Integrated Math Reasoning

Ruotao Xu, Yixin Ji, Yu Luo, Jinpeng Li, Dong Li, Peifeng Li, Juntao Li, Min Zhang
arXiv: 2604.08281v1 发布: 2026-04-09 更新: 2026-04-09

AI 摘要

论文提出ATTC框架,通过置信度自适应地决定是否信任工具结果,提升工具集成推理模型性能。

主要贡献

  • 发现并定义了工具集成推理中“工具忽略”问题
  • 提出了自适应工具信任校准(ATTC)框架
  • 实验证明ATTC能有效提升模型性能

方法论

引入ATTC框架,基于代码块置信度来决定信任或忽略工具结果,并在多个数据集上进行实验验证。

原文摘要

Large reasoning models (LRMs) have achieved strong performance enhancement through scaling test time computation, but due to the inherent limitations of the underlying language models, they still have shortcomings in tasks that require precise computation and extensive knowledge reserves. Tool-Integrated Reasoning (TIR) has emerged as a promising paradigm that incorporates tool call and execution within the reasoning trajectory. Although recent works have released some powerful open-source TIR models, our analysis reveals that these models still suffer from critical deficiencies. We find that when the reasoning of the model conflicts with the tool results, the model tends to believe in its own reasoning. And there are cases where the tool results are correct but are ignored by the model, resulting in incorrect answers, which we define as "Tool Ignored''. This indicates that the model does not know when to trust or ignore the tool. To overcome these limitations, We introduce Adaptive Tool Trust Calibration (ATTC), a novel framework that guides the model to adaptively choose to trust or ignore the tool results based on the confidence score of generated code blocks. The experimental results from various open-source TIR models of different sizes and across multiple datasets demonstrate that ATTC effectively reduces the "Tool Ignored" issue, resulting in a performance increase of 4.1% to 7.5%.

标签

LLM 工具集成 推理 信任校准 代码生成

arXiv 分类

cs.CL