What Models Know, How Well They Know It: Knowledge-Weighted Fine-Tuning for Learning When to Say "I Don't Know"
AI 摘要
论文提出了一种知识加权微调方法,使模型能够区分已知和未知问题并表达不确定性。
主要贡献
- 提出知识加权微调方法,解决知识对齐问题
- 使用多样本推理估计细粒度的知识得分
- 鼓励模型对超出范围的查询返回"I don't know"
方法论
通过多样本推理估计知识得分,并根据得分调整学习信号,鼓励模型表达不确定性。
原文摘要
While large language models (LLMs) demonstrate strong capabilities across diverse user queries, they still suffer from hallucinations, often arising from knowledge misalignment between pre-training and fine-tuning. To address this misalignment, we reliably estimate a fine-grained, instance-level knowledge score via multi-sampled inference. Using the knowledge score, we scale the learning signal according to the model's existing knowledge, while encouraging explicit "I don't know" responses for out-of-scope queries. Experimental results show that this approach allows the model to explicitly express uncertainty when it lacks knowledge, while maintaining accuracy on questions it can answer. Furthermore, we propose evaluation metrics for uncertainty, showing that accurate discrimination between known and unknown instances consistently improves performance.