UNIGEOCLIP: Unified Geospatial Contrastive Learning
AI 摘要
UNIGEOCLIP提出了一种统一的多模态对比学习框架,用于对齐地理空间数据,提升跨模态检索和推理能力。
主要贡献
- 提出了UNIGEOCLIP框架,用于联合对齐五种地理空间模态
- 采用全对全对比对齐方式,实现跨模态比较、检索和推理
- 提出了缩放的经纬度编码器,以捕获多尺度地理结构
方法论
UNIGEOCLIP利用对比学习,将航空图像、街景视图、高程模型、文本和地理坐标对齐到统一的嵌入空间,使用全对全对比对齐和改进的地理编码器。
原文摘要
The growing availability of co-located geospatial data spanning aerial imagery, street-level views, elevation models, text, and geographic coordinates offers a unique opportunity for multimodal representation learning. We introduce UNIGEOCLIP, a massively multimodal contrastive framework to jointly align five complementary geospatial modalities in a single unified embedding space. Unlike prior approaches that fuse modalities or rely on a central pivot representation, our method performs all-to-all contrastive alignment, enabling seamless comparison, retrieval, and reasoning across arbitrary combinations of modalities. We further propose a scaled latitude-longitude encoder that improves spatial representation by capturing multi-scale geographic structure. Extensive experiments across downstream geospatial tasks demonstrate that UNIGEOCLIP consistently outperforms single-modality contrastive models and coordinate-only baselines, highlighting the benefits of holistic multimodal geospatial alignment. A reference implementation is available at https://gastruc.github.io/unigeoclip.