MirrorBench: Evaluating Self-centric Intelligence in MLLMs by Introducing a Mirror
AI 摘要
论文提出MirrorBench基准,评估多模态大模型在自参照理解方面的能力,揭示其与人类的差距。
主要贡献
- 提出MirrorBench基准,用于评估MLLM的自中心智能
- 基于心理学镜像自识别测试(MSR)
- 实验表明现有MLLM在自参照理解方面存在局限性
方法论
设计基于仿真的分层任务框架,模拟镜像环境,评估MLLM从视觉感知到高级自我表征的能力。
原文摘要
Recent progress in Multimodal Large Language Models (MLLMs) has demonstrated remarkable advances in perception and reasoning, suggesting their potential for embodied intelligence. While recent studies have evaluated embodied MLLMs in interactive settings, current benchmarks mainly target capabilities to perceive, understand, and interact with external objects, lacking a systematic evaluation of self-centric intelligence. To address this, we introduce MirrorBench, a simulation-based benchmark inspired by the classical Mirror Self-Recognition (MSR) test in psychology. MirrorBench extends this paradigm to embodied MLLMs through a tiered framework of progressively challenging tasks, assessing agents from basic visual perception to high-level self-representation. Experiments on leading MLLMs show that even at the lowest level, their performance remains substantially inferior to human performance, revealing fundamental limitations in self-referential understanding. Our study bridges psychological paradigms and embodied intelligence, offering a principled framework for evaluating the emergence of general intelligence in large models. Project page: https://fflahm.github.io/mirror-bench-page/.