AI Agents 相关度: 9/10

ARGOS: Who, Where, and When in Agentic Multi-Camera Person Search

Myungchul Kim, Kwanyong Park, Junmo Kim, In So Kweon
arXiv: 2604.12762v1 发布: 2026-04-14 更新: 2026-04-14

AI 摘要

ARGOS是一个多摄像头行人搜索的交互式推理基准和框架,聚焦于时空推理和工具利用。

主要贡献

  • 提出了新的交互式多摄像头行人搜索基准ARGOS
  • 构建了包含时空拓扑图(STTG)的推理框架
  • 证明了现有LLM在该任务上的不足,以及领域特定工具的重要性

方法论

构建包含摄像头连接和过渡时间的时空拓扑图,智能体根据模糊信息进行提问、规划和排除,使用空间和时间工具辅助推理。

原文摘要

We introduce ARGOS, the first benchmark and framework that reformulates multi-camera person search as an interactive reasoning problem requiring an agent to plan, question, and eliminate candidates under information asymmetry. An ARGOS agent receives a vague witness statement and must decide what to ask, when to invoke spatial or temporal tools, and how to interpret ambiguous responses, all within a limited turn budget. Reasoning is grounded in a Spatio-Temporal Topology Graph (STTG) encoding camera connectivity and empirically validated transition times. The benchmark comprises 2,691 tasks across 14 real-world scenarios in three progressive tracks: semantic perception (Who), spatial reasoning (Where), and temporal reasoning (When). Experiments with four LLM backbones show the benchmark is far from solved (best TWS: 0.383 on Track 2, 0.590 on Track 3), and ablations confirm that removing domain-specific tools drops accuracy by up to 49.6 percentage points.

标签

AI Agents Multimodal Learning Reasoning Benchmarking

arXiv 分类

cs.CV cs.AI cs.MA