Exploring Temporal Representation in Neural Processes for Multimodal Action Prediction
AI 摘要
论文改进了基于条件神经过程的多模态动作预测模型,提升了时间信息的表征能力。
主要贡献
- 提出 DMBN-PTE 模型,改进了时间信息的表征
- 发现了 DMBN 模型在泛化性上的不足
- 在机器人自监督多模态动作预测任务中应用条件神经过程
方法论
基于条件神经过程(CNP)构建深度模态融合网络(DMBN),并引入位置时间编码(PTE)增强时间信息表征,进行自监督学习。
原文摘要
Inspired by the human ability to understand and predict others, we study the applicability of Conditional Neural Processes (CNP) to the task of self-supervised multimodal action prediction in robotics. Following recent results regarding the ontogeny of the Mirror Neuron System (MNS), we focus on the preliminary objective of self-actions prediction. We find a good MNS-inspired model in the existing Deep Modality Blending Network (DMBN), able to reconstruct the visuo-motor sensory signal during a partially observed action sequence by leveraging the probabilistic generation of CNP. After a qualitative and quantitative evaluation, we highlight its difficulties in generalizing to unseen action sequences, and identify the cause in its inner representation of time. Therefore, we propose a revised version, termed DMBN-Positional Time Encoding (DMBN-PTE), that facilitates learning a more robust representation of temporal information, and provide preliminary results of its effectiveness in expanding the applicability of the architecture. DMBN-PTE figures as a first step in the development of robotic systems that autonomously learn to forecast actions on longer time scales refining their predictions with incoming observations.