The Impact of Steering Large Language Models with Persona Vectors in Educational Applications
AI 摘要
研究表明,在教育应用中,基于激活的Persona向量会降低LLM生成答案的质量,且影响因任务和模型架构而异。
主要贡献
- 首次系统性地研究了激活控制的Persona特征在教育生成和评分中的影响。
- 揭示了Persona引导对答案质量的负面影响,特别是在开放式ELA提示上。
- 观察到评分者Persona与评分校准之间的可预测关系,以及任务和模型架构对这种关系的影响。
方法论
在ASAP-SAS基准上,通过Persona向量对七种性格特征进行短答案生成和自动评分的实验研究,并分析结果。
原文摘要
Activation-based steering can personalize large language models at inference time, but its effects in educational settings remain unclear. We study persona vectors for seven character traits in short-answer generation and automated scoring on the ASAP-SAS benchmark across three models spanning two architectures. Persona steering lowers answer quality overall, with much larger effects on open-ended English Language Arts (ELA) prompts than on factual science prompts; interpretive and argumentative tasks are up to 11x more sensitive. On the scoring side, we observe predictable valence-aligned calibration shifts: evil and impolite scorers grade more harshly, while good and optimistic scorers grade more leniently. ELA tasks are 2.5-3x more susceptible to scorer personalization than science tasks, and the Mixture-of-Experts model shows roughly 6x larger calibration shifts than the dense models. To our knowledge, this is the first study to systematically examine the effects of activation-steered persona traits in educational generation and scoring, and the results highlight the need for task-aware and architecture-aware calibration when deploying steered models in educational settings.