AI Agents 相关度: 10/10

UniToolCall: Unifying Tool-Use Representation, Data, and Evaluation for LLM Agents

Yijuan Liang, Xinghao Chen, Yifan Ge, Ziyi Wu, Hao Wu, Changyu Zeng, Wei Xing, Xiaoyu Shen
arXiv: 2604.11557v1 发布: 2026-04-13 更新: 2026-04-13

AI 摘要

UniToolCall统一了LLM智能体工具使用的数据、表示和评估,提升了工具使用性能。

主要贡献

  • 提出了UniToolCall统一框架,标准化工具学习流程。
  • 构建了包含22k+工具和390k+实例的混合训练语料库。
  • 引入Anchor Linkage机制,增强多轮推理的连贯性。

方法论

构建统一的工具学习框架,结合标准化数据集和结构化控制的合成轨迹,通过Anchor Linkage增强多轮推理。

原文摘要

Tool-use capability is a fundamental component of LLM agents, enabling them to interact with external systems through structured function calls. However, existing research exhibits inconsistent interaction representations, largely overlooks the structural distribution of tool-use trajectories, and relies on incompatible evaluation benchmarks. We present UniToolCall, a unified framework for tool learning that standardizes the entire pipeline from toolset construction and dataset generation to evaluation. The framework curates a large tool pool of 22k+ tools and constructs a hybrid training corpus of 390k+ instances by combining 10 standardized public datasets with structurally controlled synthetic trajectories. It explicitly models diverse interaction patterns, including single-hop vs. multi-hop and single-turn vs. multi-turn, while capturing both serial and parallel execution structures. To support coherent multi-turn reasoning, we further introduce an Anchor Linkage mechanism that enforces cross-turn dependencies. Furthermore, we convert 7 public benchmarks into a unified Query--Action--Observation--Answer (QAOA) representation with fine-grained evaluation at the function-call, turn, and conversation levels. Experiments show that fine-tuning Qwen3-8B on our dataset substantially improves tool-use performance. Under the distractor-heavy Hybrid-20 setting, achieves 93.0% single-turn Strict Precision, outperforming commercial models including GPT, Gemini, and Claude.

标签

LLM Agent Tool-Use Benchmark Function Call

arXiv 分类

cs.AI