Multimodal Learning 相关度: 9/10

PokeGym: A Visually-Driven Long-Horizon Benchmark for Vision-Language Models

Ruizhi Zhang, Ye Huang, Yuangang Pan, Chuanfu Shen, Zhilin Liu, Ting Xie, Wen Li, Lixin Duan
arXiv: 2604.08340v1 发布: 2026-04-09 更新: 2026-04-09

AI 摘要

PokeGym是一个视觉驱动的长程视觉语言模型基准,旨在评估3D具身环境中的VLM性能。

主要贡献

  • 提出了 PokeGym 基准,一个基于 Pokemon Legends: Z-A 的 3D 开放世界环境。
  • PokeGym 采用代码级隔离,确保纯粹基于视觉的决策。
  • 揭示了现有 VLM 在物理死锁恢复方面的局限性,并分析了死锁类型。

方法论

设计了包含导航、交互和混合场景的30个任务,并通过三种指令粒度评估VLM的视觉理解、语义推理和自主探索能力。

原文摘要

While Vision-Language Models (VLMs) have achieved remarkable progress in static visual understanding, their deployment in complex 3D embodied environments remains severely limited. Existing benchmarks suffer from four critical deficiencies: (1) passive perception tasks circumvent interactive dynamics; (2) simplified 2D environments fail to assess depth perception; (3) privileged state leakage bypasses genuine visual processing; and (4) human evaluation is prohibitively expensive and unscalable. We introduce PokeGym, a visually-driven long-horizon benchmark instantiated within Pokemon Legends: Z-A, a visually complex 3D open-world Role-Playing Game. PokeGym enforces strict code-level isolation: agents operate solely on raw RGB observations while an independent evaluator verifies success via memory scanning, ensuring pure vision-based decision-making and automated, scalable assessment. The benchmark comprises 30 tasks (30-220 steps) spanning navigation, interaction, and mixed scenarios, with three instruction granularities (Visual-Guided, Step-Guided, Goal-Only) to systematically deconstruct visual grounding, semantic reasoning, and autonomous exploration capabilities. Our evaluation reveals a key limitation of current VLMs: physical deadlock recovery, rather than high-level planning, constitutes the primary bottleneck, with deadlocks showing a strong negative correlation with task success. Furthermore, we uncover a metacognitive divergence: weaker models predominantly suffer from Unaware Deadlocks (oblivious to entrapment), whereas advanced models exhibit Aware Deadlocks (recognizing entrapment yet failing to recover). These findings highlight the need to integrate explicit spatial intuition into VLM architectures. The code and benchmark will be available on GitHub.

标签

VLM Benchmark 3D Environment Embodied AI Long-Horizon

arXiv 分类

cs.CV cs.AI