Multimodal Learning 相关度: 9/10

E-VLA: Event-Augmented Vision-Language-Action Model for Dark and Blurred Scenes

Jiajun Zhai, Hao Shi, Shangwei Guo, Kailun Yang, Kaiwei Wang
arXiv: 2604.04834v1 发布: 2026-04-06 更新: 2026-04-06

AI 摘要

E-VLA通过事件相机增强VLA模型,提升其在光线不足和模糊场景下的操作鲁棒性。

主要贡献

  • 提出E-VLA框架,利用事件流改善恶劣环境下的感知
  • 构建了包含RGB-event-action的真实世界数据集
  • 提出了轻量级的事件融合策略

方法论

使用事件相机获取事件流,通过事件融合(如叠加)或事件适配器,增强VLA模型在图像质量差时的感知能力。

原文摘要

Robotic Vision-Language-Action (VLA) models generalize well for open-ended manipulation, but their perception is fragile under sensing-stage degradations such as extreme low light, motion blur, and black clipping. We present E-VLA, an event-augmented VLA framework that improves manipulation robustness when conventional frame-based vision becomes unreliable. Instead of reconstructing images from events, E-VLA directly leverages motion and structural cues in event streams to preserve semantic perception and perception-action consistency under adverse conditions. We build an open-source teleoperation platform with a DAVIS346 event camera and collect a real-world synchronized RGB-event-action manipulation dataset across diverse tasks and illumination settings. We also propose lightweight, pretrained-compatible event integration strategies and study event windowing and fusion for stable deployment. Experiments show that even a simple parameter-free fusion, i.e., overlaying accumulated event maps onto RGB images, could substantially improve robustness in dark and blur-heavy scenes: on Pick-Place at 20 lux, success increases from 0% (image-only) to 60% with overlay fusion and to 90% with our event adapter; under severe motion blur (1000 ms exposure), Pick-Place improves from 0% to 20-25%, and Sorting from 5% to 32.5%. Overall, E-VLA provides systematic evidence that event-driven perception can be effectively integrated into VLA models, pointing toward robust embodied intelligence beyond conventional frame-based imaging. Code and dataset will be available at https://github.com/JJayzee/E-VLA.

标签

事件相机 视觉语言行动模型 机器人操作 鲁棒性 多模态

arXiv 分类

cs.CV cs.MM cs.RO eess.IV