Multimodal Learning 相关度: 9/10

Can Vision Language Models Judge Action Quality? An Empirical Evaluation

Miguel Monte e Freitas, Rui Henriques, Ricardo Rei, Pedro Henrique Martins
arXiv: 2604.08294v1 发布: 2026-04-09 更新: 2026-04-09

AI 摘要

评估视觉语言模型在动作质量评估中的性能,发现现有模型表现不佳,存在偏差。

主要贡献

  • 全面评估VLM在动作质量评估中的性能
  • 揭示VLM在AQA任务中的系统性偏差
  • 建立了VLM在AQA领域的研究基准

方法论

使用多个VLM模型,在不同动作领域、任务、表示和提示策略下进行评估,并分析预测分布。

原文摘要

Action Quality Assessment (AQA) has broad applications in physical therapy, sports coaching, and competitive judging. Although Vision Language Models (VLMs) hold considerable promise for AQA, their actual performance in this domain remains largely uncharacterised. We present a comprehensive evaluation of state-of-the-art VLMs across activity domains (e.g. fitness, figure skating, diving), tasks, representations, and prompting strategies. Baseline results reveal that Gemini 3.1 Pro, Qwen3-VL and InternVL3.5 models perform only marginally above random chance, and although strategies such as incorporation of skeleton information, grounding instructions, reasoning structures and in-context learning lead to isolated gains, none is consistently effective. Analysis of prediction distributions uncovers two systematic biases: a tendency to predict correct execution regardless of visual evidence, and a sensitivity to superficial linguistic framing. Reformulating tasks contrastively to mitigate these biases yields minimal improvement, suggesting that the models' limitations go beyond these biases, pointing to a fundamental difficulty with fine-grained movement quality assessment. Our findings establish a rigorous baseline for future VLM-based AQA research and provide an actionable outline for failure modes requiring mitigation prior to reliable real-world deployment.

标签

VLM 动作质量评估 多模态学习 偏差分析

arXiv 分类

cs.CV cs.AI cs.CL