LLM Reasoning 相关度: 9/10

Less Approximates More: Harmonizing Performance and Confidence Faithfulness via Hybrid Post-Training for High-Stakes Tasks

Haokai Ma, Lee Yan Zhen, Gang Yang, Yunshan Ma, Ee-Chien Chang, Tat-Seng Chua
arXiv: 2604.08454v1 发布: 2026-04-09 更新: 2026-04-09

AI 摘要

论文提出HyTuning框架,通过渐进推理增益自适应融合推理蒸馏和强化学习,提升LLM在高风险任务中的准确性和置信度可靠性。

主要贡献

  • 提出渐进推理增益(PRG)衡量推理步骤对最终答案的支持度
  • 提出HyTuning框架,自适应融合推理蒸馏(RD)和强化学习(RLIF)
  • 实验证明HyTuning在有限监督下提升LLM准确性和置信度可靠性

方法论

HyTuning利用PRG指标,结合少量监督推理轨迹和大量无标签查询,自适应调整RD和RLIF的权重。

原文摘要

Large language models are increasingly deployed in high-stakes tasks, where confident yet incorrect inferences may cause severe real-world harm, bringing the previously overlooked issue of confidence faithfulness back to the forefront. A promising solution is to jointly optimize unsupervised Reinforcement Learning from Internal Feedback (RLIF) with reasoning-trace-guided Reasoning Distillation (RD), which may face three persistent challenges: scarcity of high-quality training corpora, factually unwarranted overconfidence and indiscriminate fusion that amplifies erroneous updates. Inspired by the human confidence accumulation from uncertainty to certainty, we propose Progressive Reasoning Gain (PRG) to measure whether reasoning steps progressively strengthen support for the final answer. Furthermore, we introduce HyTuning, a hybrid post-training framework that adaptively reweights RD and RLIF via a PRG-style metric, using scarce supervised reasoning traces as a stable anchor while exploiting abundant unlabeled queries for scalability. Experiments on several domain-specific and general benchmarks demonstrate that HyTuning improves accuracy while achieving confidence faithfulness under limited supervision, supporting a practical "Less Approximates More" effect.

标签

LLM Reasoning Confidence Calibration Reinforcement Learning Distillation

arXiv 分类

cs.LG