What Do Vision-Language Models Encode for Personalized Image Aesthetics Assessment?
AI 摘要
分析视觉语言模型(VLM)在个性化图像美学评估(PIAA)中的表征能力,并用于轻量级个性化。
主要贡献
- 分析了VLMs内部的美学属性表征
- 利用VLMs的表征进行轻量级PIAA
- 分析了不同VLM架构和图像域之间的美学信息传递
方法论
通过分析VLM内部表征来研究其编码的美学属性,并利用线性模型进行PIAA,避免了微调。
原文摘要
Personalized image aesthetics assessment (PIAA) is an important research problem with practical real-world applications. While methods based on vision-language models (VLMs) are promising candidates for PIAA, it remains unclear whether they internally encode rich, multi-level aesthetic attributes required for effective personalization. In this paper, we first analyze the internal representations of VLMs to examine the presence and distribution of such aesthetic attributes, and then leverage them for lightweight, individual-level personalization without model fine-tuning. Our analysis reveals that VLMs encode diverse aesthetic attributes that propagate into the language decoder layers. Building on these representations, we demonstrate that simple linear models can perform PIAA effectively. We further analyze how aesthetic information is transferred across layers in different VLM architectures and across image domains. Our findings provide insights into how VLMs can be utilized for modeling subjective, individual aesthetic preferences. Our code is available at https://github.com/ynklab/vlm-latent-piaa.