LLM Reasoning 相关度: 9/10

General365: Benchmarking General Reasoning in Large Language Models Across Diverse and Challenging Tasks

Junlin Liu, Shengnan An, Shuang Zhou, Dan Ma, Shixiong Luo, Ying Xie, Yuan Zhang, Wenling Yuan, Yifan Zhou, Xiaoyu Li, Ziwen Wang, Xuezhi Cao, Xunliang Cai
arXiv: 2604.11778v1 发布: 2026-04-13 更新: 2026-04-13

AI 摘要

提出了General365基准测试,评估LLM在通用推理任务上的能力,发现其表现远低于领域特定任务。

主要贡献

  • 构建了General365通用推理基准,包含365个种子问题和1095个变体问题
  • 揭示了当前LLM在通用推理能力上的不足
  • 提供了代码、数据集和排行榜,促进相关研究

方法论

通过设计K-12水平知识背景下的推理问题,评估26个LLM模型在General365基准上的表现,分析其通用推理能力。

原文摘要

Contemporary large language models (LLMs) have demonstrated remarkable reasoning capabilities, particularly in specialized domains like mathematics and physics. However, their ability to generalize these reasoning skills to more general and broader contexts--often termed general reasoning--remains under-explored. Unlike domain-specific reasoning, general reasoning relies less on expert knowledge but still presents formidable reasoning challenges, such as complex constraints, nested logical branches, and semantic interference. To address this gap, we introduce General365, a benchmark specifically designed to assess general reasoning in LLMs. By restricting background knowledge to a K-12 level, General365 explicitly decouples reasoning from specialized expertise. The benchmark comprises 365 seed problems and 1,095 variant problems across eight categories, ensuring both high difficulty and diversity. Evaluations across 26 leading LLMs reveal that even the top-performing model achieves only 62.8% accuracy, in stark contrast to the near-perfect performances of LLMs in math and physics benchmarks. These results suggest that the reasoning abilities of current LLMs are heavily domain-dependent, leaving significant room for improvement in broader applications. We envision General365 as a catalyst for advancing LLM reasoning beyond domain-specific tasks toward robust, general-purpose real-world scenarios. Code, Dataset, and Leaderboard: https://general365.github.io

标签

LLM 推理 基准测试 通用推理

arXiv 分类

cs.CL cs.AI