AI Agents 相关度: 9/10

From Reactive to Proactive: Assessing the Proactivity of Voice Agents via ProVoice-Bench

Ke Xu, Yuhao Wang, Yu Wang
arXiv: 2604.15037v1 发布: 2026-04-16 更新: 2026-04-16

AI 摘要

提出了ProVoice-Bench,一个评估主动语音代理的框架,揭示了现有模型在主动性方面的不足。

主要贡献

  • 设计了首个主动语音代理评估框架ProVoice-Bench
  • 构建了包含1182个高质量样本的数据集
  • 揭示了现有多模态LLM在主动性方面的局限性

方法论

通过多阶段数据合成流程,构建数据集,并利用该数据集评估现有Multimodal LLM的性能表现。

原文摘要

Recent advancements in LLM agents are gradually shifting from reactive, text-based paradigms toward proactive, multimodal interaction. However, existing benchmarks primarily focus on reactive responses, overlooking the complexities of proactive intervention and monitoring. To bridge this gap, we introduce ProVoice-Bench, the first evaluation framework specifically designed for proactive voice agents, featuring four novel tasks. By leveraging a multi-stage data synthesis pipeline, we curate 1,182 high-quality samples for rigorous testing. Our evaluation of state-of-the-art Multimodal LLMs reveals a significant performance gap, particularly regarding over-triggering and reasoning capabilities. These findings highlight the limitations of current models and offer a roadmap for developing more natural, context-aware proactive agents.

标签

语音代理 主动性 多模态 评估框架

arXiv 分类

cs.AI cs.CL cs.SD