Controlling Distributional Bias in Multi-Round LLM Generation via KL-Optimized Fine-Tuning
AI 摘要
提出一种基于KL散度和Kahneman-Tversky优化的微调框架,用于控制LLM生成结果的分布偏差。
主要贡献
- 提出了一种新的微调框架,耦合steering token校准和语义对齐。
- 引入混合目标函数,结合KL散度和Kahneman-Tversky优化。
- 实验证明该方法在属性生成任务中优于基线方法,实现了精确的分布控制。
方法论
通过KL散度锚定潜在steering token的概率质量,并用Kahneman-Tversky优化将token与语义一致的响应绑定,进行模型微调。
原文摘要
While the real world is inherently stochastic, Large Language Models (LLMs) are predominantly evaluated on single-round inference against fixed ground truths. In this work, we shift the lens to distribution alignment: assessing whether LLMs, when prompted repeatedly, can generate outputs that adhere to a desired target distribution, e.g. reflecting real-world statistics or a uniform distribution. We formulate distribution alignment using the attributes of gender, race, and sentiment within occupational contexts. Our empirical analysis reveals that off-the-shelf LLMs and standard alignment techniques, including prompt engineering and Direct Preference Optimization, fail to reliably control output distributions. To bridge this gap, we propose a novel fine-tuning framework that couples Steering Token Calibration with Semantic Alignment. We introduce a hybrid objective function combining Kullback-Leibler divergence to anchor the probability mass of latent steering tokens and Kahneman-Tversky Optimization to bind these tokens to semantically consistent responses. Experiments across six diverse datasets demonstrate that our approach significantly outperforms baselines, achieving precise distributional control in attribute generation tasks.