AI Agents 相关度: 8/10

Saliency-Guided Representation with Consistency Policy Learning for Visual Unsupervised Reinforcement Learning

Jingbo Sun, Qichao Zhang, Songjun Tu, Xing Fang, Yupeng Zheng, Haoran Li, Ke Chen, Dongbin Zhao
arXiv: 2604.05931v1 发布: 2026-04-07 更新: 2026-04-07

AI 摘要

提出SRCP框架,通过显著性引导表示和一致性策略学习,提升视觉无监督强化学习的泛化能力。

主要贡献

  • 提出显著性引导的动态任务学习表征
  • 设计快速采样一致性策略
  • 在视觉无监督强化学习中实现SOTA零样本泛化

方法论

SRCP解耦表征学习和后继训练,通过显著性引导的动态任务学习表征,并结合一致性策略学习来建模技能条件策略。

原文摘要

Zero-shot unsupervised reinforcement learning (URL) offers a promising direction for building generalist agents capable of generalizing to unseen tasks without additional supervision. Among existing approaches, successor representations (SR) have emerged as a prominent paradigm due to their effectiveness in structured, low-dimensional settings. However, SR methods struggle to scale to high-dimensional visual environments. Through empirical analysis, we identify two key limitations of SR in visual URL: (1) SR objectives often lead to suboptimal representations that attend to dynamics-irrelevant regions, resulting in inaccurate successor measures and degraded task generalization; and (2) these flawed representations hinder SR policies from modeling multi-modal skill-conditioned action distributions and ensuring skill controllability. To address these limitations, we propose Saliency-Guided Representation with Consistency Policy Learning (SRCP), a novel framework that improves zero-shot generalization of SR methods in visual URL. SRCP decouples representation learning from successor training by introducing a saliency-guided dynamics task to capture dynamics-relevant representations, thereby improving successor measure and task generalization. Moreover, it integrates a fast-sampling consistency policy with URL-specific classifier-free guidance and tailored training objectives to improve skill-conditioned policy modeling and controllability. Extensive experiments on 16 tasks across 4 datasets from the ExORL benchmark demonstrate that SRCP achieves state-of-the-art zero-shot generalization in visual URL and is compatible with various SR methods.

标签

无监督强化学习 后继表示 零样本泛化 显著性引导

arXiv 分类

cs.CV cs.AI