How LLMs Follow Instructions: Skillful Coordination, Not a Universal Mechanism
AI 摘要
指令遵循并非通用机制,而是语言模型协调多种语言能力的技能。
主要贡献
- 揭示指令遵循并非单一机制,而是多技能协调
- 通过诊断探针分析,验证指令遵循的技能组合特性
- 分析了任务复杂性与模型层级之间的关系
方法论
通过对三个指令微调模型在九个任务上进行诊断探测,包括探针训练、跨任务迁移和因果消融分析。
原文摘要
Instruction tuning is commonly assumed to endow language models with a domain-general ability to follow instructions, yet the underlying mechanism remains poorly understood. Does instruction-following rely on a universal mechanism or compositional skill deployment? We investigate this through diagnostic probing across nine diverse tasks in three instruction-tuned models. Our analysis provides converging evidence against a universal mechanism. First, general probes trained across all tasks consistently underperform task-specific specialists, indicating limited representational sharing. Second, cross-task transfer is weak and clustered by skill similarity. Third, causal ablation reveals sparse asymmetric dependencies rather than shared representations. Tasks also stratify by complexity across layers, with structural constraints emerging early and semantic tasks emerging late. Finally, temporal analysis shows constraint satisfaction operates as dynamic monitoring during generation rather than pre-generation planning. These findings indicate that instruction-following is better characterized as skillful coordination of diverse linguistic capabilities rather than deployment of a single abstract constraint-checking process.