EpiBench: Benchmarking Multi-turn Research Workflows for Multimodal Agents
AI 摘要
EpiBench是一个多轮多模态的科研工作流基准测试,评估智能体在科研任务中的表现。
主要贡献
- 提出了EpiBench基准测试
- 引入了过程级别的评估框架
- 揭示了现有模型在复杂科研任务中的不足
方法论
构建包含文献搜索、数据对齐、证据整合等环节的科研任务,评估智能体在多轮交互中的表现。
原文摘要
Scientific research follows multi-turn, multi-step workflows that require proactively searching the literature, consulting figures and tables, and integrating evidence across papers to align experimental settings and support reproducible conclusions. This joint capability is not systematically assessed in existing benchmarks, which largely under-evaluate proactive search, multi-evidence integration and sustained evidence use over time. In this work, we introduce EpiBench, an episodic multi-turn multimodal benchmark that instantiates short research workflows. Given a research task, agents must navigate across papers over multiple turns, align evidence from figures and tables, and use the accumulated evidence in the memory to answer objective questions that require cross paper comparisons and multi-figure integration. EpiBench introduces a process-level evaluation framework for fine-grained testing and diagnosis of research agents. Our experiments show that even the leading model achieves an accuracy of only 29.23% on the hard split, indicating substantial room for improvement in multi-turn, multi-evidence research workflows, providing an evaluation platform for verifiable and reproducible research agents.