Multimodal Learning 相关度: 8/10

OV-Stitcher: A Global Context-Aware Framework for Training-Free Open-Vocabulary Semantic Segmentation

Seungjae Moon, Seunghyun Oh, Youngmin Ro
arXiv: 2604.08110v1 发布: 2026-04-09 更新: 2026-04-09

AI 摘要

OV-Stitcher通过拼接子图像特征,实现全局上下文感知的免训练开放词汇语义分割。

主要贡献

  • 提出OV-Stitcher框架,用于解决免训练开放词汇语义分割中的全局上下文问题
  • 通过在最终编码器块中拼接碎片化的子图像特征,实现全局注意力
  • 在多个基准测试中取得了显著的mIoU提升,验证了框架的有效性

方法论

使用滑动窗口策略处理子图像,并在最终编码器块中重建注意力表示,实现全局上下文聚合,生成语义对齐的分割图。

原文摘要

Training-free open-vocabulary semantic segmentation(TF-OVSS) has recently attracted attention for its ability to perform dense prediction by leveraging the pretrained knowledge of large vision and vision-language models, without requiring additional training. However, due to the limited input resolution of these pretrained encoders, existing TF-OVSS methods commonly adopt a sliding-window strategy that processes cropped sub-images independently. While effective for managing high-resolution inputs, this approach prevents global attention over the full image, leading to fragmented feature representations and limited contextual reasoning. We propose OV-Stitcher, a training-free framework that addresses this limitation by stitching fragmented sub-image features directly within the final encoder block. By reconstructing attention representations from fragmented sub-image features, OV-Stitcher enables global attention within the final encoder block, producing coherent context aggregation and spatially consistent, semantically aligned segmentation maps. Extensive evaluations across eight benchmarks demonstrate that OV-Stitcher establishes a scalable and effective solution for open-vocabulary segmentation, achieving a notable improvement in mean Intersection over Union(mIoU) from 48.7 to 50.7 compared with prior training-free baselines.

标签

语义分割 开放词汇 视觉语言模型 全局上下文

arXiv 分类

cs.CV cs.AI cs.LG