Multimodal Learning 相关度: 9/10

MLLM-as-a-Judge Exhibits Model Preference Bias

Shuitsu Koyama, Yuiga Wada, Daichi Yashima, Komei Sugiura
arXiv: 2604.11589v1 发布: 2026-04-13 更新: 2026-04-13

AI 摘要

研究发现多模态大语言模型评估存在模型偏好,提出 Philautia-Eval 并设计 Pomms 模型缓解偏见。

主要贡献

  • 提出 Philautia-Eval 方法量化模型偏好
  • 发现 MLLM-as-a-Judge 存在自偏好和家族偏好
  • 提出 Pomms 模型有效缓解模型偏好

方法论

设计 Philautia-Eval 解耦偏好倾向与生成质量差异,通过大量实验分析不同 MLLM 的偏好,并提出集成方法。

原文摘要

Automatic evaluation using multimodal large language models (MLLMs), commonly referred to as MLLM-as-a-Judge, has been widely used to measure model performance. If such MLLM-as-a-Judge methods were biased, they could distort model comparisons and benchmark-driven scientific progress. However, it remains unclear to what extent MLLM-as-a-Judge methods favor or disfavor text generated by specific MLLMs. In this study, we propose Philautia-Eval to investigate such model-specific preference bias. Philautia-Eval quantifies the degree of the bias by disentangling preference tendencies from differences in generation quality. Using 1.29M caption-score pairs collected from 12 MLLMs, we found that representative MLLMs tend to exhibit self-preference bias. Moreover, experimental results indicate mutual preference bias within particular model families, which is potentially driven by reused connectors and overlapping instruction-tuning resources. Finally, we introduce a simple ensemble of MLLMs, Pomms. Our results demonstrated that Pomms effectively mitigated the model-specific preference bias while maintaining performance.

标签

MLLM Bias Evaluation Multimodal Learning Model Preference

arXiv 分类

cs.CV