Multimodal Learning 相关度: 9/10

KITE: Keyframe-Indexed Tokenized Evidence for VLM-Based Robot Failure Analysis

Mehdi Hosseinzadeh, King Hang Wong, Feras Dayoub
arXiv: 2604.07034v1 发布: 2026-04-08 更新: 2026-04-08

AI 摘要

KITE利用关键帧索引和符号化证据,提升VLM在机器人故障分析中的性能。

主要贡献

  • 提出了一种训练无关的机器人故障分析前端KITE
  • 将长视频转化为紧凑、可解释的符号化证据
  • 在RoboFAC基准测试上显著提升了VLM的故障检测和定位能力

方法论

KITE提取运动关键帧,结合BEV图和符号化信息,构成统一提示,用于驱动VLM进行故障分析。

原文摘要

We present KITE, a training-free, keyframe-anchored, layout-grounded front-end that converts long robot-execution videos into compact, interpretable tokenized evidence for vision-language models (VLMs). KITE distills each trajectory into a small set of motion-salient keyframes with open-vocabulary detections and pairs each keyframe with a schematic bird's-eye-view (BEV) representation that encodes relative object layout, axes, timestamps, and detection confidence. These visual cues are serialized with robot-profile and scene-context tokens into a unified prompt, allowing the same front-end to support failure detection, identification, localization, explanation, and correction with an off-the-shelf VLM. On the RoboFAC benchmark, KITE with Qwen2.5-VL substantially improves over vanilla Qwen2.5-VL in the training-free setting, with especially large gains on simulation failure detection, identification, and localization, while remaining competitive with a RoboFAC-tuned baseline. A small QLoRA fine-tune further improves explanation and correction quality. We also report qualitative results on real dual-arm robots, demonstrating the practical applicability of KITE as a structured and interpretable front-end for robot failure analysis. Code and models are released on our project page: https://m80hz.github.io/kite/

标签

机器人 故障分析 视觉语言模型 关键帧

arXiv 分类

cs.RO cs.AI cs.CV