DetailVerifyBench: A Benchmark for Dense Hallucination Localization in Long Image Captions
AI 摘要
DetailVerifyBench是一个用于长图像描述中密集幻觉定位的基准测试,包含高质量图像和token级别的幻觉标注。
主要贡献
- 提出了DetailVerifyBench基准测试,用于评估长图像描述中幻觉的精确定位能力。
- 该基准测试包含1000张图像,覆盖五个不同的领域,具有高多样性。
- 提供了token级别的幻觉标注,使得能够进行细粒度的幻觉检测评估。
方法论
构建包含多样化图像和长描述的数据集,并对描述中的幻觉进行token级别的标注,从而形成一个评估基准。
原文摘要
Accurately detecting and localizing hallucinations is a critical task for ensuring high reliability of image captions. In the era of Multimodal Large Language Models (MLLMs), captions have evolved from brief sentences into comprehensive narratives, often spanning hundreds of words. This shift exponentially increases the challenge: models must now pinpoint specific erroneous spans or words within extensive contexts, rather than merely flag response-level inconsistencies. However, existing benchmarks lack the fine granularity and domain diversity required to evaluate this capability. To bridge this gap, we introduce DetailVerifyBench, a rigorous benchmark comprising 1,000 high-quality images across five distinct domains. With an average caption length of over 200 words and dense, token-level annotations of multiple hallucination types, it stands as the most challenging benchmark for precise hallucination localization in the field of long image captioning to date. Our benchmark is available at https://zyx-hhnkh.github.io/DetailVerifyBench/.