Visual Prompt Based Reasoning for Offroad Mapping using Multimodal LLMs
AI 摘要
该论文提出了一种基于视觉提示和多模态LLM的零样本越野环境可行驶区域推理方法。
主要贡献
- 提出了基于SAM2分割和VLM推理的零样本越野环境可行驶区域识别方法
- 无需训练特定地形模型,依赖VLM的推理能力
- 在模拟环境中实现了超越现有模型的导航性能
方法论
利用SAM2进行图像分割,并结合数值标签传递给VLM,提示其推理可行驶区域。
原文摘要
Traditional approaches to off-road autonomy rely on separate models for terrain classification, height estimation, and quantifying slip or slope conditions. Utilizing several models requires training each component separately, having task specific datasets, and fine-tuning. In this work, we present a zero-shot approach leveraging SAM2 for environment segmentation and a vision-language model (VLM) to reason about drivable areas. Our approach involves passing to the VLM both the original image and the segmented image annotated with numeric labels for each mask. The VLM is then prompted to identify which regions, represented by these numeric labels, are drivable. Combined with planning and control modules, this unified framework eliminates the need for explicit terrain-specific models and relies instead on the inherent reasoning capabilities of the VLM. Our approach surpasses state-of-the-art trainable models on high resolution segmentation datasets and enables full stack navigation in our Isaac Sim offroad environment.