LLM Reasoning 相关度: 9/10

Self-Debias: Self-correcting for Debiasing Large Language Models

Xuan Feng, Shuai Zhao, Luwei Xiao, Tianlong Gu, Bo An
arXiv: 2604.08243v1 发布: 2026-04-09 更新: 2026-04-09

AI 摘要

提出Self-Debias框架,通过动态约束和在线自提升机制,提升LLM的去偏能力。

主要贡献

  • 提出Self-Debias框架,实现LLM的自我纠偏。
  • 使用细粒度的轨迹级别目标函数,动态约束偏见。
  • 集成在线自提升机制,利用一致性过滤生成监督信号。

方法论

将去偏过程视为资源重新分配问题,通过动态去偏约束调整模型输出概率,并使用在线自提升机制。

原文摘要

Although Large Language Models (LLMs) demonstrate remarkable reasoning capabilities, inherent social biases often cascade throughout the Chain-of-Thought (CoT) process, leading to continuous "Bias Propagation". Existing debiasing methods primarily focus on static constraints or external interventions, failing to identify and interrupt this propagation once triggered. To address this limitation, we introduce Self-Debias, a progressive framework designed to instill intrinsic self-correction capabilities. Specifically, we reformulate the debiasing process as a strategic resource redistribution problem, treating the model's output probability mass as a limited resource to be reallocated from biased heuristics to unbiased reasoning paths. Unlike standard preference optimization which applies broad penalties, Self-Debias employs a fine-grained trajectory-level objective subject to dynamic debiasing constraints. This enables the model to selectively revise biased reasoning suffixes while preserving valid contextual prefixes. Furthermore, we integrate an online self-improvement mechanism utilizing consistency filtering to autonomously synthesize supervision signals. With merely 20k annotated samples, Self-Debias activates efficient self-correction, achieving superior debiasing performance while preserving general reasoning capabilities without continuous external oversight.

标签

LLM Debiasing Self-Correction Reasoning Chain-of-Thought

arXiv 分类

cs.CL