LLM Reasoning 相关度: 9/10

SUPERNOVA: Eliciting General Reasoning in LLMs with Reinforcement Learning on Natural Instructions

Ashima Suvarna, Kendrick Phan, Mehrab Beikzadeh, Hritik Bansal, Saadia Gabriel
arXiv: 2604.08477v1 发布: 2026-04-09 更新: 2026-04-09

AI 摘要

SUPERNOVA提出一个数据策展框架,利用强化学习提升LLM的通用推理能力,并验证了数据设计对性能的影响。

主要贡献

  • 提出SUPERNOVA框架,增强RLVR在通用推理上的表现
  • 系统性分析了数据选择、混合策略和数据增强对推理性能的影响
  • 实验表明SUPERNOVA在多个推理基准测试上优于现有模型

方法论

通过强化学习,利用专家标注的指令微调数据集,系统性地调整数据以增强通用推理能力,并进行大量实验验证。

原文摘要

Reinforcement Learning with Verifiable Rewards (RLVR) has significantly improved large language model (LLM) reasoning in formal domains such as mathematics and code. Despite these advancements, LLMs still struggle with general reasoning tasks requiring capabilities such as causal inference and temporal understanding. Extending RLVR to general reasoning is fundamentally constrained by the lack of high-quality, verifiable training data that spans diverse reasoning skills. To address this challenge, we propose SUPERNOVA, a data curation framework for RLVR aimed at enhancing general reasoning. Our key insight is that instruction-tuning datasets containing expert-annotated ground-truth encode rich reasoning patterns that can be systematically adapted for RLVR. To study this, we conduct 100+ controlled RL experiments to analyze how data design choices impact downstream reasoning performance. In particular, we investigate three key factors: (i) source task selection, (ii) task mixing strategies, and (iii) synthetic interventions for improving data quality. Our analysis reveals that source task selection is non-trivial and has a significant impact on downstream reasoning performance. Moreover, selecting tasks based on their performance for individual target tasks outperforms strategies based on overall average performance. Finally, models trained on SUPERNOVA outperform strong baselines (e.g., Qwen3.5) on challenging reasoning benchmarks including BBEH, Zebralogic, and MMLU-Pro. In particular, training on SUPERNOVA yields relative improvements of up to 52.8\% on BBEH across model sizes, demonstrating the effectiveness of principled data curation for RLVR. Our findings provide practical insights for curating human-annotated resources to extend RLVR to general reasoning. The code and data is available at https://github.com/asuvarna31/supernova.

标签

强化学习 通用推理 数据策展 指令微调

arXiv 分类

cs.AI cs.LG