Multimodal Learning 相关度: 9/10

Multi-Modal Sensor Fusion using Hybrid Attention for Autonomous Driving

Mayank Mayank, Bharanidhar Duraisamy, Florian Geiß, Abhinav Valada
arXiv: 2604.04797v1 发布: 2026-04-06 更新: 2026-04-06

AI 摘要

提出MMF-BEV,一种基于混合注意力机制的雷达-相机BEV融合框架,用于自动驾驶3D目标检测。

主要贡献

  • 提出MMF-BEV融合框架
  • 引入Deformable Attention进行跨模态特征对齐
  • 提出两阶段训练策略

方法论

构建相机和雷达BEV分支,利用Deformable Self-Attention增强特征,通过Deformable Cross-Attention进行融合。

原文摘要

Accurate 3D object detection for autonomous driving requires complementary sensors. Cameras provide dense semantics but unreliable depth, while millimeter-wave radar offers precise range and velocity measurements with sparse geometry. We propose MMF-BEV, a radar-camera BEV fusion framework that leverages deformable attention for cross-modal feature alignment on the View-of-Delft (VoD) 4D radar dataset [1]. MMF-BEV builds a BEVDepth [2] camera branch and a RadarBEVNet [3] radar branch, each enhanced with Deformable Self-Attention, and fuses them via a Deformable Cross-Attention module. We evaluate three configurations: camera-only, radar-only, and hybrid fusion. A sensor contribution analysis quantifies per-distance modality weighting, providing interpretable evidence of sensor complementarity. A two-stage training strategy - pre-training the camera branch with depth supervision, then jointly training radar and fusion modules stabilizes learning. Experiments on VoD show that MMF-BEV consistently outperforms unimodal baselines and achieves competitive results against prior fusion methods across all object classes in both the full annotated area and near-range Region of Interest.

标签

多模态融合 自动驾驶 3D目标检测 BEV 注意力机制

arXiv 分类

cs.CV cs.LG