SeLaR: Selective Latent Reasoning in Large Language Models
AI 摘要
SeLaR提出了一种轻量级无训练框架,通过选择性潜在推理提升LLM的推理能力。
主要贡献
- 提出了熵门控机制,选择性地激活低置信步骤的软嵌入。
- 提出了熵感知对比正则化,鼓励探索多个潜在推理路径。
- 在五个推理基准测试中,SeLaR优于标准CoT和现有无训练方法。
方法论
SeLaR利用熵值来控制软嵌入的激活,并使用对比正则化来避免软嵌入的坍塌,从而提升推理性能。
原文摘要
Chain-of-Thought (CoT) has become a cornerstone of reasoning in large language models, yet its effectiveness is constrained by the limited expressiveness of discrete token sampling. Recent latent reasoning approaches attempt to alleviate this limitation by replacing discrete tokens with soft embeddings (probability-weighted mixtures of token embeddings) or hidden states, but they commonly suffer from two issues: (1) global activation injects perturbations into high-confidence steps, impairing reasoning stability; and (2) soft embeddings quickly collapse toward the highest-probability token, limiting exploration of alternative trajectories. To address these challenges, we propose SeLaR (Selective Latent Reasoning), a lightweight and training-free framework. SeLaR introduces an entropy-gated mechanism that activates soft embeddings only at low-confidence steps, while preserving discrete decoding at high-confidence steps. Additionally, we propose an entropy-aware contrastive regularization that pushes soft embeddings away from the dominant (highest-probability) token's direction, encouraging sustained exploration of multiple latent reasoning paths. Experiments on five reasoning benchmarks demonstrate that SeLaR consistently outperforms standard CoT and state-of-the-art training-free methods.