AI Agents 相关度: 9/10

HINTBench: Horizon-agent Intrinsic Non-attack Trajectory Benchmark

Jiacheng Wang, Jinchang Hou, Fabian Wang, Ping Jian, Chenfu Bao, Zhonghou Lv
arXiv: 2604.13954v1 发布: 2026-04-15 更新: 2026-04-15

AI 摘要

HINTBench:评估智能体在良性条件下内在风险的新基准测试。

主要贡献

  • 提出了非攻击性内在风险审计的概念
  • 构建了包含629条智能体轨迹的HINTBench基准
  • 分析了现有模型在内在风险检测和诊断方面的不足

方法论

构建包含风险和安全轨迹的数据集,定义风险类型分类,并评估现有模型在风险检测、定位和诊断方面的性能。

原文摘要

Existing agent-safety evaluation has focused mainly on externally induced risks. Yet agents may still enter unsafe trajectories under benign conditions. We study this complementary but underexplored setting through the lens of \emph{intrinsic} risk, where intrinsic failures remain latent, propagate across long-horizon execution, and eventually lead to high-consequence outcomes. To evaluate this setting, we introduce \emph{non-attack intrinsic risk auditing} and present \textbf{HINTBench}, a benchmark of 629 agent trajectories (523 risky, 106 safe; 33 steps on average) supporting three tasks: risk detection, risk-step localization, and intrinsic failure-type identification. Its annotations are organized under a unified five-constraint taxonomy. Experiments reveal a substantial capability gap: strong LLMs perform well on trajectory-level risk detection, but their performance drops to below 35 Strict-F1 on risk-step localization, while fine-grained failure diagnosis proves even harder. Existing guard models transfer poorly to this setting. These findings establish intrinsic risk auditing as an open challenge for agent safety.

标签

智能体安全 内在风险 基准测试 风险评估

arXiv 分类

cs.LG cs.AI