CRFT: Consistent-Recurrent Feature Flow Transformer for Cross-Modal Image Registration
AI 摘要
CRFT提出了一种基于特征流学习的跨模态图像配准统一框架,实现了高精度和鲁棒性。
主要贡献
- 提出了一致循环特征流Transformer (CRFT) 框架
- 引入了迭代差异引导注意力机制和空间几何变换 (SGT)
- 在跨模态图像配准任务上超越了现有方法
方法论
CRFT使用Transformer架构学习模态无关的特征流表示,通过粗到精的方式进行特征对齐和光流估计,并使用迭代注意力机制优化。
原文摘要
We present Consistent-Recurrent Feature Flow Transformer (CRFT), a unified coarse-to-fine framework based on feature flow learning for robust cross-modal image registration. CRFT learns a modality-independent feature flow representation within a transformer-based architecture that jointly performs feature alignment and flow estimation. The coarse stage establishes global correspondences through multi-scale feature correlation, while the fine stage refines local details via hierarchical feature fusion and adaptive spatial reasoning. To enhance geometric adaptability, an iterative discrepancy-guided attention mechanism with a Spatial Geometric Transform (SGT) recurrently refines the flow field, progressively capturing subtle spatial inconsistencies and enforcing feature-level consistency. This design enables accurate alignment under large affine and scale variations while maintaining structural coherence across modalities. Extensive experiments on diverse cross-modal datasets demonstrate that CRFT consistently outperforms state-of-the-art registration methods in both accuracy and robustness. Beyond registration, CRFT provides a generalizable paradigm for multimodal spatial correspondence, offering broad applicability to remote sensing, autonomous navigation, and medical imaging. Code and datasets are publicly available at https://github.com/NEU-Liuxuecong/CRFT.